Loading image...Kiro

Product

  • About Kiro
  • IDE
  • CLI
  • Web
  • Mobile
  • Crew
  • Pricing
  • Downloads

For

  • Enterprise
  • Startups
  • Students

Community

  • Overview
  • Ambassadors
  • Discord
  • Events
  • Powers
  • Shop
  • Showcase

Resources

  • Docs
  • Blog
  • Changelog
  • FAQs
  • Report a bug
  • Suggest an idea
  • Billing support

Social

Site TermsLicenseResponsible AI PolicyLegalPrivacy PolicyCookie Preferences
Loading image...Kiro
  • Enterprise
  • Pricing
  • Docs
SIGN INDOWNLOADS
Loading image...Kiro

Get Started

InstallationAuthenticationYour first project

Models

OverviewAvailable modelsReasoning effort

Features

How Kiro works
Specs
Steering
Hooks
MCP
Permissions
Custom agents
Agent Skills
Powers
CompactionKiroignoreCheckpoints and rewind
Built-in tools
Configuration scopes

IDE 1.x

What's new in 1.0
Setup & First Run
Editor
Chat
Experimental
Troubleshooting0.x reference

CLI

What's new in 3.0
Setup & First Run
Terminal UI
Chat
Headless modeACPAuto complete
Experimental
2.x reference

Crew

Quick startInstallationRunning 24/7
Chat
Agent Capabilities
Features
Interfaces
Apps
ConfigurationSecurityTroubleshooting

Web - Preview

Setup & First RunIdentity Center
Connect your repositories
Working with the agent
Autonomous modeAutomations
Sandbox

Mobile - Preview

Overview

Commands and Reference

CLI commandsSlash commandsBuilt-in toolsExit codesSettings

Billing

OverviewManaging your subscriptionUpgrading your planDowngrading your planCancelling your planPurchasing add-on creditsManaging your paymentsManaging usage notificationsManaging your taxesContacting billing supportDeleting your accountRelated questions

Enterprise

ConceptsOnboarding quickstart
Connecting your identity provider
Subscribe your teamManage subscriptions
Governance
Monitor and track
SettingsManaged updatesBillingIAMSupported regions

Privacy and Security

OverviewData protectionCode referencesCompliance validationInfrastructure securityIAM permissionsFirewalls, proxies, and data perimetersVPC endpoints (AWS PrivateLink)

Guides

Overview
Language support
Learn by playing

Migration

Migrating from Q DeveloperMigrating from VSCodeUpgrading from Q CLI

Models


Kiro gives you access to frontier and open weight AI models from OpenAI, Anthropic, and other providers. GPT-5.6 brings OpenAI models to Kiro for the first time, with three tiers that balance agentic performance and cost. Pick the right model for the job, or select Auto to let Kiro route each task to the optimal model automatically.

Info

Model selection is only available for chat experience - interactive and non-interactive.

Quick comparison

ModelContextCostRegionsFreeProPro+Pro MaxPower
GPT-5.6 Sol272K2.4xUS, EU✓✓✓✓
GPT-5.6 Terra272K1.0xUS, EU✓✓✓✓
GPT-5.6 Luna272K0.1xUS, EU✓✓✓✓
Claude Opus 51M2.2xUS, EU✓✓✓✓
Claude Opus 4.81M2.2xUS, EU✓✓✓✓
Claude Opus 4.71M2.2xUS, EU✓✓✓✓
Claude Opus 4.61M2.2xUS, EU✓✓✓✓
Claude Opus 4.5200K2.2xUS, EU✓✓✓✓
Claude Sonnet 51M1.3xUS, EU✓✓✓✓
Claude Sonnet 4.61M1.3xUS, EU✓✓✓✓
Claude Sonnet 4.5200K1.3xUS, EU✓✓✓✓✓
Claude Sonnet 4.0200K1.3xUS, EU✓✓✓✓✓
Auto—1.0xUS, EU✓✓✓✓✓
Claude Haiku 4.5200K0.4xUS, EU✓✓✓✓
DeepSeek 3.2128K0.25xUS, EU✓✓✓✓✓
MiniMax M2.5200K0.25xUS, EU✓✓✓✓✓
GLM-5200K0.5xUS, EU✓✓✓✓✓
MiniMax M2.1200K0.15xUS, EU✓✓✓✓✓
Qwen3 Coder Next256K0.05xUS, EU✓✓✓✓✓

US and EU are geographies, each spanning several AWS Regions rather than a single Region. All authentication methods are supported for every model. For the Regions in each geography and the endpoint that serves your requests, see Inference endpoint regions.

Cost is relative to Auto (1.0x baseline). For example, a task that costs 10 credits on Auto would cost 22 credits on Opus, 4 credits on Haiku, or 0.5 credits on Qwen3 Coder Next.

Info

Models that share the same credit multiplier won't necessarily consume the same number of credits per task. Actual consumption depends on factors like how many tokens the model generates, internal thinking depth, and tokenizer differences. For example, Opus 4.8 uses an updated tokenizer compared to Opus 4.6, so the same prompt and response can be counted as a different number of tokens - leading to different credit costs even though both carry a 2.2x multiplier.

Inference endpoint regions

The geography that serves your request depends on both the model you select and the region of your Kiro profile. GPT-5.6 models are served from the US regardless of your profile region. Every other model is served from the geography that matches your profile.

Profile region applies to enterprise users who sign in through IAM Identity Center or an external identity provider. Free Tier users and individual subscribers are always served from the US.

ModelKiro profile in US East (N. Virginia)Kiro profile in Europe (Frankfurt)
GPT-5.6 SolUSUS
GPT-5.6 TerraUSUS
GPT-5.6 LunaUSUS
Claude Opus 5USEU
Claude Opus 4.8USEU
Claude Opus 4.7USEU
Claude Opus 4.6USEU
Claude Opus 4.5USEU
Claude Sonnet 5USEU
Claude Sonnet 4.6USEU
Claude Sonnet 4.5USEU
Claude Sonnet 4.0USEU
AutoUSEU
Claude Haiku 4.5USEU
DeepSeek 3.2USEU
MiniMax M2.5USEU
GLM-5USEU
MiniMax M2.1USEU
Qwen3 Coder NextUSEU

Regions in each geography

Kiro is powered by Amazon Bedrock, which uses cross-region inference to distribute requests across the Regions within a geography. Your request can be processed in any Region listed for your geography.

GeographyRegions used for inference
USUS East (N. Virginia) us-east-1, US West (Oregon) us-west-2, US East (Ohio) us-east-2, AWS GovCloud (US-East), AWS GovCloud (US-West)
EUEurope (Frankfurt) eu-central-1, Europe (Ireland) eu-west-1, Europe (Paris) eu-west-3, Europe (Stockholm) eu-north-1, Europe (Milan) eu-south-1, Europe (Spain) eu-south-2

Cross-region inference does not change where your data is stored. Models marked experimental are an exception to the table above: they may be processed in commercial AWS Regions worldwide, including outside your profile's geography. See data protection for the full reference, and Amazon Bedrock cross-Region inference for how inference profiles route requests.

How to switch models

Use the model dropdown in the chat interface to switch models. Your selection applies to all subsequent messages in the conversation.

Which model should you use?

Use caseModelWhy
General developmentAutoRoutes to the optimal model per task, balances quality and cost automatically
Hardest multi-step developmentGPT-5.6 SolBest fit for long-horizon refactors and terminal work that requires sustained planning and tool coordination
Routine multi-step developmentGPT-5.6 TerraBalanced tier for everyday agentic work; its 1.0x Kiro credit multiplier sits between Luna's 0.1x and Sol's 2.4x
High-frequency agentic workGPT-5.6 LunaFastest, lowest-cost GPT-5.6 tier for repeated tasks where throughput matters; 0.1x Kiro credit multiplier
Highest reliabilityOpus 5State-of-the-art on agentic coding benchmarks, strongest multi-agent coordination, completes full tasks rather than leaving stubs
Near-Opus agentic at lower costSonnet 5Approaches Opus 4.8 on reasoning and tool use, plans before editing, runs longer autonomously
Speed or credit savingsHaiku 4.5Near-frontier intelligence at a fraction of the cost, great for quick iterations and sub-agents
Frontier coding at low costMiniMax M2.5Near Opus-level results at 0.25x cost, strong across the full development lifecycle
Repo-scale agentic workGLM-5200K context optimized for long-horizon workflows across large codebases
Long coding sessions on a budgetQwen3 Coder Next256K context with strong error recovery at 0.05x cost

Model availability

Model availability can vary by country or region. Kiro's model offerings align with each provider's usage and geographic requirements. For more information, see supported countries and regions: OpenAI, Anthropic, MiniMax, Zhipu AI (GLM), DeepSeek, and Qwen.

Reasoning effort

For models that support configurable reasoning effort, you can control how much reasoning the model applies to your prompts. Lower effort levels produce faster, shorter responses and use fewer credits. Higher levels spend more tokens on deeper analysis, multi-step reasoning, and thorough code generation.

CapabilityIDECLIWebMobile
Reasoning effort selection✓✓——

Set the level from the model selector's Effort panel in the IDE, or with /effort (or the --effort launch flag) in the CLI. Your choice persists, and the picker only shows levels supported by your current model. See Reasoning effort for the full reference: per-surface mechanics, supported models, persistent per-model defaults, thinking behavior, and precedence.

Info

Higher effort levels use more tokens internally, which means more credit consumption per interaction - even at the same credit multiplier. This is one reason why two models with identical multipliers can produce different credit costs for the same task.

Best practices

  • Start with Auto for most work. It optimizes both quality and cost automatically.
  • Use GPT-5.6 Sol for the hardest multi-step work - long-horizon refactors and complex terminal tasks where you need state-of-the-art agentic performance.
  • Use GPT-5.6 Terra or Luna when you want strong agentic capability with throughput or cost as the primary concern. Terra for balanced performance, Luna for maximum efficiency.
  • Switch to Opus 5 when you hit a wall on a complex problem, need sustained multi-file work, or want the strongest multi-agent coordination and code review accuracy.
  • Use Sonnet 5 when you want strong agentic behavior at a lower cost than Opus, especially for multi-step tasks that need to run to completion.
  • Use Haiku for quick iterations, simple fixes, or when you want to conserve credits.
  • Monitor your usage in your account settings to understand how model choice affects consumption.
  • Factor model cost into your tier: If you primarily use Opus, consider Pro+, Pro Max, or Power for more credits. See plans and billing for details.

For detailed descriptions of each model's capabilities, strengths, and lifecycle status, see Available models.

Page updated: August 4, 2026
Your first project
Available models