Loading image...Kiro

Product

  • About Kiro
  • Agents
  • IDE
  • CLI
  • Web
  • Mobile
  • Crew
  • Pricing
  • Downloads

For

  • Enterprise
  • Startups
  • Students

Community

  • Overview
  • Ambassadors
  • Case studies
  • Discord
  • Events
  • Powers
  • Shop
  • Showcase

Resources

  • Docs
  • Blog
  • Changelog
  • FAQs
  • Report a bug
  • Suggest an idea
  • Billing support

Social

Site TermsLicenseResponsible AI PolicyLegalPrivacy PolicyCookie Preferences
Loading image...Kiro
  • Agents
  • Enterprise
  • Pricing
  • Docs
SIGN INDOWNLOADS
Loading image...Kiro

Get Started

InstallationAuthenticationYour first project

Models

OverviewAvailable modelsReasoning effortAWS GovCloud (US) models

Features

How Kiro worksACP integrations
Specs
Steering
Hooks
MCP
Permissions
Custom agents
Workflows
Agent Skills
Powers
Cloud sessionsCompactionKiroignoreCheckpoints and rewind
Built-in tools
Configuration scopes

IDE 1.x

What's new in 1.0
Setup & First Run
Editor
Chat
Experimental
Troubleshooting0.x reference

CLI

What's new in V3
Setup & First Run
Terminal UI
Chat
Fullscreen modeVoice modeHeadless modeACPAuto complete
Experimental
2.x reference

Crew

Quick startInstallationRunning 24/7
Chat
Agent Capabilities
Features
Interfaces
Apps
System & storageConfigurationSecurityTroubleshooting

Web

Setup & First RunIdentity Center
Connect your repositories
Working with the agent
Autonomous modeAutomationsMemoryConfiguration Sync
Sandbox

Mobile - Preview

Overview

Commands and Reference

CLI commandsSlash commandsBuilt-in toolsExit codesSettings

Billing

OverviewManaging your subscriptionUpgrading your planDowngrading your planCancelling your planPurchasing add-on creditsManaging your paymentsManaging usage notificationsManaging your taxesContacting billing supportDeleting your accountRelated questions

Enterprise

ConceptsOnboarding quickstart
Connecting your identity provider
Deployment optionsSubscribe your teamManage subscriptions
Governance
Monitor and track
SettingsManaged updatesBillingIAMSupported regions

Privacy and Security

OverviewData protectionCode referencesCompliance validationInfrastructure securityIAM permissionsFirewalls, proxies, and data perimetersVPC endpoints (AWS PrivateLink)

Guides

Overview
Language support
Learn by playing

Migration

Migrating from Q DeveloperMigrating from VSCodeUpgrading from Q CLI
View as Markdown

Models

View as Markdown

Kiro gives you access to frontier and open weight AI models from OpenAI, Anthropic, and other providers. GPT-5.6 brings OpenAI models to Kiro for the first time, with three tiers that balance agentic performance and cost. Pick the right model for the job, or select Auto to let Kiro route each task to the optimal model automatically.

Info

Model selection is only available for chat experience - interactive and non-interactive.

Quick comparison

ModelContextCostRegionsFreeProPro+Pro MaxPower
GPT-5.6 Sol1M4.4x*US, EU✓✓✓✓
GPT-5.6 Terra1M2.2x*US, EU✓✓✓✓
GPT-5.6 Luna1M0.6x*US, EU✓✓✓✓
Claude Fable 5.1†1M6xUS East only
Claude Opus 5.51M2.0xUS, EU✓✓✓✓
Claude Sonnet 5.51M1.3xUS, EU✓✓✓✓
Claude Opus 51M2.2xUS, EU✓✓✓✓
Claude Opus 4.81M2.2xUS, EU✓✓✓✓
Claude Opus 4.71M2.2xUS, EU✓✓✓✓
Claude Opus 4.61M2.2xUS, EU✓✓✓✓
Claude Opus 4.5200K2.2xUS, EU✓✓✓✓
Claude Sonnet 51M1.3xUS, EU✓✓✓✓
Claude Sonnet 4.61M1.3xUS, EU✓✓✓✓
Claude Sonnet 4.5200K1.3xUS, EU✓✓✓✓✓
Claude Sonnet 4.0200K1.3xUS, EU✓✓✓✓✓
Auto—1.0xUS, EU✓✓✓✓✓
Claude Haiku 4.5200K0.4xUS, EU✓✓✓✓
DeepSeek 3.2128K0.25xUS, EU✓✓✓✓✓
MiniMax M2.5200K0.25xUS, EU✓✓✓✓✓
GLM-5200K0.5xUS, EU✓✓✓✓✓
MiniMax M2.1200K0.15xUS, EU✓✓✓✓✓
Qwen3 Coder Next256K0.05xUS, EU✓✓✓✓✓

† Claude Fable 5.1 is rolling out in a limited Preview for Kiro Enterprise customers. Administrators enable access through model governance.

Enterprise Preview rollout

Claude Fable 5.1 is rolling out gradually to Kiro Enterprise organizations with model governance enabled. It has a 1M context window and a 6x credit multiplier. Inference runs only in US East (N. Virginia) (us-east-1). Organizations with access can add Fable 5.1 from Manage approved list and make it available to their users. See Claude Fable 5.1 for capabilities and data-handling details.

US and EU are commercial geographies, each spanning several AWS Regions rather than a single Region. For the Regions in each geography and the endpoint that serves your requests, see Inference endpoint regions.

Cost is relative to Auto (1.0x baseline). For example, a task that costs 10 credits on Auto would cost 20 credits on Opus 5.5, 22 credits on Opus 5, 4 credits on Haiku, or 0.5 credits on Qwen3 Coder Next.

* GPT-5.6 uses OpenAI and Amazon Bedrock's two-tier context pricing model: requests up to 272K tokens use the listed short context rate, and requests over 272K tokens are billed at double that rate (Sol 8.8x, Terra 4.4x, Luna 1.2x).

Info

Models that share the same credit multiplier won't necessarily consume the same number of credits per task. Actual consumption depends on factors like how many tokens the model generates, internal thinking depth, and tokenizer differences. For example, Opus 4.8 uses an updated tokenizer compared to Opus 4.6, so the same prompt and response can be counted as a different number of tokens - leading to different credit costs even though both carry a 2.2x multiplier.

AWS GovCloud (US)

AWS GovCloud (US) has a separate model catalog and Region-specific availability. See AWS GovCloud (US) models for the current model list, East and West availability, access requirements, and model-authorization guidance.

Inference endpoint regions

The geography that serves requests from commercial profiles depends on both the model and the Kiro profile Region. GPT-5.6 models are served from the US regardless of profile Region, Claude Fable 5.1 is available only in US East (N. Virginia), and every other model is served from the geography that matches the profile.

Profile Region applies to enterprise users who sign in through IAM Identity Center or an external identity provider. Free Tier users and individual subscribers are always served from the commercial US geography.

ModelKiro profile in US East (N. Virginia)Kiro profile in Europe (Frankfurt)
GPT-5.6 SolUSUS
GPT-5.6 TerraUSUS
GPT-5.6 LunaUSUS
Claude Fable 5.1US—
Claude Opus 5.5USEU
Claude Sonnet 5.5USEU
Claude Opus 5USEU
Claude Opus 4.8USEU
Claude Opus 4.7USEU
Claude Opus 4.6USEU
Claude Opus 4.5USEU
Claude Sonnet 5USEU
Claude Sonnet 4.6USEU
Claude Sonnet 4.5USEU
Claude Sonnet 4.0USEU
AutoUSEU
Claude Haiku 4.5USEU
DeepSeek 3.2USEU
MiniMax M2.5USEU
GLM-5USEU
MiniMax M2.1USEU
Qwen3 Coder NextUSEU

Regions in each geography

Kiro is powered by Amazon Bedrock, which uses cross-region inference to distribute requests across the Regions within a geography. Your request can be processed in any Region listed for your geography.

GeographyRegions used for inference
USUS East (N. Virginia) us-east-1, US West (Oregon) us-west-2, US East (Ohio) us-east-2
EUEurope (Frankfurt) eu-central-1, Europe (Ireland) eu-west-1, Europe (Paris) eu-west-3, Europe (Stockholm) eu-north-1, Europe (Milan) eu-south-1, Europe (Spain) eu-south-2

Cross-region inference does not change where your data is stored. For commercial profiles, models marked experimental are an exception to the table above: they may be processed in commercial AWS Regions worldwide, including outside your profile's geography. This exception does not apply to AWS GovCloud (US). See data protection for the full reference, Kiro in AWS GovCloud (US) for GovCloud behavior, and Amazon Bedrock cross-Region inference for how inference profiles route requests.

How to switch models

Use the model dropdown in the chat interface to switch models. Your selection applies to all subsequent messages in the conversation.

Which model should you use?

Use caseModelWhy
General developmentAutoRoutes to the optimal model per task, balances quality and cost automatically
Hardest multi-step developmentGPT-5.6 SolBest fit for long-horizon refactors and terminal work that requires sustained planning and tool coordination; 1M context for longer codebases, documents, and multi-turn histories
Routine multi-step developmentGPT-5.6 TerraBalanced tier for everyday agentic work; its 2.2x Kiro credit multiplier applies to requests up to 272K tokens
High-frequency agentic workGPT-5.6 LunaFastest, lowest-cost GPT-5.6 tier for repeated tasks where throughput matters; its 0.6x Kiro credit multiplier applies to requests up to 272K tokens
Demanding agentic codingOpus 5.5Strong fit for long, complex tasks that benefit from sustained tool use, self-verification, and clear communication
Fast scoped coding and agentsSonnet 5.5Clear upgrade over Sonnet 5 with more than 30% faster output; strong fit for bug fixes, review, and scoped docs, slides, and spreadsheets
Speed or credit savingsHaiku 4.5Near-frontier intelligence at a fraction of the cost, great for quick iterations and sub-agents
Frontier coding at low costMiniMax M2.5Near Opus-level results at 0.25x cost, strong across the full development lifecycle
Repo-scale agentic workGLM-5200K context optimized for long-horizon workflows across large codebases
Long coding sessions on a budgetQwen3 Coder Next256K context with strong error recovery at 0.05x cost

Model availability

Model availability can vary by country or region. Kiro's model offerings align with each provider's usage and geographic requirements. For more information, see supported countries and regions: OpenAI, Anthropic, MiniMax, Zhipu AI (GLM), DeepSeek, and Qwen.

Reasoning effort

For models that support configurable reasoning effort, you can control how much reasoning the model applies to your prompts. Lower effort levels produce faster, shorter responses and use fewer credits. Higher levels spend more tokens on deeper analysis, multi-step reasoning, and thorough code generation.

CapabilityIDECLIWebMobile
Reasoning effort selection✓✓✓—

In the IDE, set the level with the Effort slider below the model list in the Model panel. In Kiro CLI, open /model and choose Effort, or use the --effort launch flag. In Kiro Web, use the reasoning effort selector. The IDE choice applies to subsequent messages in the conversation. In CLI, effort chosen through /model persists for that model. Direct /effort <level> changes only the current session unless you then run /effort set-current-as-default; in V3, --effort is also session-only. Web choices apply to the current session. Each picker shows only levels supported by your current model. See Reasoning effort for the full reference: per-surface mechanics, supported models, persistent per-model defaults, thinking behavior, and precedence.

Info

Higher effort levels use more tokens internally, which means more credit consumption per interaction - even at the same credit multiplier. This is one reason why two models with identical multipliers can produce different credit costs for the same task.

Best practices

  • Start with Auto for most work. It optimizes both quality and cost automatically.
  • Use GPT-5.6 Sol for the hardest multi-step work - long-horizon refactors and complex terminal tasks where you need state-of-the-art agentic performance.
  • Use GPT-5.6 Terra or Luna when you want strong agentic capability with throughput or cost as the primary concern. Terra for balanced performance, Luna for maximum efficiency.
  • Use Opus 5.5 when you hit a wall on a complex problem, need sustained multi-file work, or want stronger self-verification and clearer communication across a long task.
  • Use Sonnet 5.5 for fast, scoped coding, bug fixes, review, agent and subagent work, or focused documentation, slides, and spreadsheets.
  • Use Haiku for quick iterations, simple fixes, or when you want to conserve credits.
  • Monitor your usage in your account settings to understand how model choice affects consumption.
  • Factor model cost into your tier: If you primarily use Opus, consider Pro+, Pro Max, or Power for more credits. See plans and billing for details.

For detailed descriptions of each model's capabilities, strengths, and lifecycle status, see Available models.

Page updated: October 8, 2026
Your first project
Available models