Meet Kimi K3, Moonshot AI's most advanced AI model yet. With 2.8 trillion parameters and a massive one-million-token context window, Kimi K3 can handle complex tasks, large codebases, long documents, and deep research in a single workflow.
Built-in AI agents work directly inside the chat, allowing Kimi to plan, reason, and complete multi-step tasks from start to finish. Whether you are coding, researching, writing, analysing information, or building autonomous agent workflows, Kimi AI provides one powerful platform for everything.
Since launch, K3's open weights have shipped on Hugging Face, K3 has joined GitHub Copilot alongside K2.7 Code, and the API gained selectable low / high / max reasoning effort. The platform also includes K2.7 Code HighSpeed, delivering coding performance at up to 260 tokens per second.
Everything that has changed since the K3 launch - open weights, GitHub Copilot, new API controls, retirements, benchmarks and the regulatory picture. Checked against Moonshot's platform docs, the Kimi-K3 GitHub repo and major press as of 10 September 2026.
Moonshot shipped the full 2.8T-parameter K3 checkpoint to Hugging Face (moonshotai/Kimi-K3) ahead of its July 27 deadline - the largest open-weight model ever released. It ships as a native MXFP4 checkpoint (MXFP8 activations, ~1.4 TB), with 104B active parameters per token (16 of 896 experts + 2 shared), and day-zero recipes for vLLM, SGLang and TokenSpeed.
The license is the new Kimi K3 License - MIT-style, but with two extra conditions: model-as-a-service businesses above $20M revenue in 12 months need a separate agreement, and products with 100M+ monthly users or $20M+ monthly revenue must show UI attribution. It is not the Modified MIT license used by the K2 family.
On August 6 GitHub added Kimi K3 to Copilot for Pro, Pro+, Max, Business and Enterprise (hosted on Fireworks AI, usage-based $3 / $0.30 cached / $15 per 1M; off by default for Business and Enterprise until an admin enables it). The API also gained selectable reasoning_effort - low, high or max (default) - thinking itself stays always-on.
NSA, FBI and CISA issued joint advisory AA26-251A accusing six China-based labs - Moonshot AI, DeepSeek, Alibaba, MiniMax, StepFun and Z.AI - of routing distillation requests through multiple pathways to extract US models. Officials had earlier alleged K3 was trained on outputs from Claude Fable 5; independent researchers questioned the 15-day gap between Fable's launch and K3's. No sanctions yet; Entity List action has been floated.
Artificial Analysis Index v4.3 places K3 at 44 (GLM-5.3: 45). Moonshot's Kimi-Vendor-Verifier now publishes a conformance leaderboard so you can check third-party hosts serve K3 faithfully.
Alongside FrontierSWE 81.2, Terminal-Bench 2.1 88.3, DeepSWE 67.5 and GDPval-AA v2 Elo 1686. See the comparison table.
The sunset announced in July is complete. Migrate to kimi-k2.7-code or kimi-k3. Weights for K2.5 remain on Hugging Face.
K3 now accepts reasoning_effort of low, high or max (default). Moonshot also updated platform rate-limit rules in August after "high-frequency abnormal requests" hit cluster stability. The 10-30% top-up bonus promo ended Aug 11.
Available in VS Code, Visual Studio, JetBrains, Xcode, Eclipse, Copilot CLI, the cloud agent, github.com and GitHub Mobile. Hosted on Fireworks AI; usage-based billing.
Moonshot closed a round at a $35B valuation, China's second most valuable private AI lab after DeepSeek. Bloomberg reported Alibaba Cloud supplied access to ~20,000 Nvidia H200s (Alibaba denied it).
Published a day early (US time). MXFP4 checkpoint, ~1.4 TB, vLLM / SGLang / TokenSpeed recipes. Hugging Face →
Seven major releases in thirteen months - each one pushing a specific capability frontier. K3 is the 2.8T-parameter open-weight flagship; the K2 family remains the workhorse lineup. K2.5's hosted API retired on August 31, 2026.
Kimi K3 is Moonshot AI's next-generation flagship model, built to evolve with the way you work. Powered by 2.8 trillion parameters (Stable LatentMoE - 896 experts, 16 active + 2 shared, 104B active per token) and Kimi Delta Attention for 6.3× faster decoding at 1M tokens, it understands long documents, analyses entire codebases, and maintains context across complex conversations.
Native text, image, and video input. Always-on thinking with selectable low / high / max effort. #1 on LMArena Frontend Code Arena (1,679 Elo), 93.5% GPQA Diamond, Terminal-Bench 2.1 88.3, SWE Marathon 42.0 - beating Claude Opus 4.8. Open weights live on Hugging Face as a native MXFP4 checkpoint under the Kimi K3 License.
The same Kimi K2.7 Code model - Moonshot's most capable K2-series coding model - served at extreme throughput. Up to 260 tokens per second on short-context tasks, ~180 tok/s on median coding inputs. Designed for agentic workflows where speed determines task completion time. Available to Kimi Code Beta, API developers, and Business users.
⚡ 260 tok/s peak · 6× faster $1.90 / $8.00 per 1MMoonshot's most capable coding model in the K2 family. Reduces reasoning token usage by ~30% vs K2.6 while improving scores on every benchmark. Mandatory thinking mode, preserve-thinking across turns. MCP tool-use SOTA: 81.1 on MCP Mark Verified, beating Claude Opus 4.8 at 76.4. Multimodal via MoonViT 400M encoder. Since July 1, 2026 it has been the first open-weight model in GitHub Copilot's model picker (Azure-hosted) - joined by K3 on August 6. Open weights under Modified MIT license.
+21.8% Kimi Code Bench v2 −30% reasoning tokens 🐙 In GitHub CopilotLong-horizon agentic coding, general availability. 300-agent swarm with 4,000 coordinated steps. Claw Groups for cross-model collaboration. Document-to-Skill conversion. 262K context window. SWE-bench Verified 80.2% - SOTA among open-source models. BrowseComp Swarm 86.3%. Supports instant mode and thinking mode - the only Kimi model with switchable thinking. The cheapest frontier model in the lineup at $0.55/$2.65 per 1M.
80.2% SWE-bench Verified 300 agents · 4,000 stepsNative multimodal intelligence trained on 15T mixed visual + text tokens. MoonViT 400M encoder for images and video. 256K context. Agent Swarm v1: 100 parallel sub-agents, 1,500 tool calls, 4.5× execution speedup. SWE-bench 76.8%, AIME 96.1%, VideoMMU 86.6%. ⚠ The hosted kimi-k2.5 API endpoint was retired on August 31, 2026 - migrate to K2.7 Code or K3. Weights remain on Hugging Face for self-hosting.
Post-trained reasoning variant with interleaved chain-of-thought and native tool use. 200–300 sequential tool calls without losing task context. Native INT4 quantization via QAT for 2× speed vs FP16 - no accuracy loss. Tencent CodeBuddy integrates K2 Thinking as its core engine. Pioneered the think→act→observe→think loop at production scale. Legacy API slugs reached EOL May 25, 2026 - self-host via open weights.
300 tool calls · 2× speed INT4The original trillion-parameter open-source frontier. 1T MoE, 32B active, 128K context, MuonClip optimizer for zero training instability across 15.5T tokens. Set the open-source agentic baseline: SWE-bench 65.8%, MMLU-Pro 73.3%, τ²-bench 80%. Modified MIT License. Still a strong cost-efficient self-hosted model for text-only workflows.
Open source · Modified MITK3's 1,048,576-token context window fits an entire 800K-line codebase in one session. Kimi Delta Attention (KDA) makes it usable - 6.3× faster decoding at 1M tokens - and the API charges no long-context surcharge: flat $3/$15 per 1M across the full window.
Up to 260 tokens per second on short-context tasks, ~180 tok/s on median coding inputs. 6× faster than standard K2.7 Code. Same model, optimized serving infrastructure. Available to Beta users.
K2.6 coordinates up to 300 sub-agents executing 4,000 steps simultaneously. K3 adds a dedicated kimi-k3-swarm-max variant for swarm workloads. Compress hours of parallel research and code generation into minutes. Available from Allegretto plan upward.
K3 and K2.7 Code always reason before responding - no shortcutting on complex tasks. Preserve-thinking keeps the chain across multi-turn sessions. New since August: K3's reasoning_effort accepts low, high or max (default), so you can trade depth for latency and cost without losing the thinking trace.
MoonViT 400M encoder processes images alongside text in K2.5, K2.6, and K2.7 Code. K3 goes further with native video input - Video-MME 90.0, MathVision 94.3 - upload screen recordings, design walkthroughs, or UI bug reproductions. Design-to-code from any visual input.
K2.7 Code scores 81.1 on MCP Mark Verified - beating Claude Opus 4.8 at 76.4 - and K3 scores 76.5 on Toolathlon-Verified. Both are in GitHub Copilot's model picker: K2.7 Code since July 1 (Azure), K3 since August 6 (Fireworks AI). No Moonshot API key needed; prompts don't route to Moonshot's servers.
Direct AI-native access to World Bank datasets, financial market data, economic indicators, and academic research - queryable in plain language inside Kimi workflows. Available from Adagio (200 requests/mo).
Migrate from GPT or Claude by changing two lines: base_url and model. Full function-calling, streaming, structured outputs, and tool use. K3 note: temperature, top_p and seed are fixed server-side (temperature 1.0, top_p 0.95) - omit them. Token-based billing separate from membership.
K2 models ship as open weights on Hugging Face under Modified MIT. K3 weights are live as a native MXFP4 checkpoint (~1.4 TB) under the Kimi K3 License - MIT-style with revenue/attribution clauses for very large deployments. Deploy with vLLM, SGLang, TokenSpeed or TensorRT-LLM. Self-hosting is the GDPR-compliant path for EU workloads - the hosted API routes via China.
Start free with Kimi K3 access, 200 Professional Data requests/month, Kimi Slides, and Deep Research - no credit card required.
Access all models, tools, and agent features. Free to start.
Start immediately, no credit card needed. You get:
The recommended starting point for regular professional use:
Token-based API billing separate from membership. Get a key at platform.kimi.ai (minimum $1 top-up; tier sets rate limits). K3: $3.00/1M input · $0.30 cached · $15.00/1M output. K2.6: $0.55/1M · $2.65/1M. K2.7 Code: $0.95/1M · $4.00/1M. K2.7 HighSpeed: $1.90/1M · $8.00/1M. K3 is also billed usage-based inside GitHub Copilot at the same rates.
Write, convert, review, and translate documents. Upload PDFs, Word files, slides, or spreadsheets - get clean structured outputs with summaries and professional formatting. Powered by K2.6 document agent. LaTeX PDF support.
Try Kimi Docs →Two creation modes: Adaptive (30–60 min, research-first with K2 Thinking) and Visual (5–10 min with Nano Banana Pro / Gemini 3 Pro Image). Chart parsing converts static chart images to native editable PPTX objects. Agentic Slides converts any document to a deck.
Try Kimi Slides →Build formulas, pivots, and dashboards from plain language. Process up to 1M rows of data in a single session. Generate charts, financial models, pivot tables, and data summaries. Export to Excel-compatible formats.
Try Kimi Sheets →Create modern, responsive websites from a prompt. Visual Coding with K3 turns design screenshots directly into production-ready code - K3 is #1 on LMArena Frontend Code Arena. Full-stack generation: frontend, backend, auth, and database ops in one session.
Try Website Builder →Terminal-first coding agent powered by K2.7 Code, with K3 available via /model k3. Works in VS Code, Cursor, JetBrains, and Zed - and both models are now in GitHub Copilot too. Autonomous coding sessions, codebase navigation, multi-file editing. Starting at $19/month. K2.7 Code HighSpeed Beta available.
Try Kimi Code →24/7 cloud-based AI agent - no server setup needed. Persistent long-term memory. 5,000+ ClawHub skills. 40GB cloud storage. Scheduled automation and 24/7 task execution. Available on Allegretto plan and above.
Try Kimi Claw →Fully OpenAI and Anthropic SDK compatible. Change two lines to switch from GPT or Claude. Token-based billing, separate from app membership. K3 live with selectable reasoning effort.
kimi-k3 - ✦ Flagship · 1M ctx · reasoning_effort low/high/max
kimi-k3-swarm-max - K3 agent swarm
kimi-k2.7-code-highspeed - 260 tok/s ⚡
kimi-k2.7-code - Coding specialist
kimi-k2.6 - General purpose GA
kimi-k2.5 / moonshot-v1 - retired Aug 31, 2026
kimi-k2-* legacy slugs - EOL May 25, 2026
Cached input: $0.19/1M (K2.x) · $0.30/1M (K3). No long-context surcharge on K3. K3 is not yet eligible for the 60% Batch API discount that K2.x models receive. Minimum top-up now $1. Get API key →
K3: thinking always on · reasoning_effort = low | high | max (default) · temperature / top_p / seed fixed server-side · images base64 or ms:// only · video via files API
K2.7 Code: thinking always on · temperature=1.0
K2.6: thinking switchable (on/off)
Access chain-of-thought via reasoning_content in response. Pass it back in multi-turn sessions to preserve thinking context. K3 default max_completion_tokens is 131,072 (max 1,048,576) - set an explicit lower limit. Rate-limit rules were tightened in August 2026.
Download from HuggingFace → moonshotai/Kimi-K3. Native MXFP4 weights / MXFP8 activations, ~1.4 TB; practical minimum is an 8-node × 8 × 80 GB cluster. Kimi K3 License (MIT-style; separate agreement above $20M MaaS revenue). Official recipes for vLLM, SGLang and TokenSpeed; TensorRT-LLM also works. Run the Kimi Vendor Verifier before production traffic - a public conformance leaderboard launched Sep 7.
Access Kimi K3 on web and mobile - no setup, no credit card. Use Agent mode for multi-step tasks; K3's thinking is always on. Free tier includes 200 Professional Data requests/month, Kimi Slides Adaptive mode, and 256K context.
Get your key at platform.kimi.ai ($1 minimum top-up). Set base_url="https://api.moonshot.ai/v1" and model="kimi-k3" in any OpenAI SDK; pick reasoning_effort low, high or max. Works with LangChain, LlamaIndex, and Anthropic SDKs. Coding default: kimi-k2.7-code.
Kimi Code CLI at kimi.com/code works in VS Code, Cursor, JetBrains, and Zed - default K2.7 Code, switch with /model k3. Or pick K3 / K2.7 Code straight from GitHub Copilot's model picker (Business/Enterprise admins must enable the K3 policy). Join kimi.com/code/beta for HighSpeed.
Download from huggingface.co/moonshotai/Kimi-K3. Native MXFP4 checkpoint (~1.4 TB). Deploy with vLLM, SGLang, TokenSpeed or TensorRT-LLM. Kimi K3 License - commercial use permitted; separate agreement only for MaaS businesses above $20M revenue.
Kimi uses music tempo-inspired plan names. All tiers include Kimi K3 access - full 1M context unlocks at Allegretto. API billing is always separate from membership.
All plans include Kimi K3 access · Full 1M context on Allegretto and above. API billing is always separate from membership. Full pricing details →
Updated September 2026 with the benchmarks Moonshot added to the K3 model card after launch, plus the first independent index scores.
Kimi K3 ✦ Moonshot AI · Jul 2026 |
Kimi K2.7 Code Moonshot AI · Jun 2026 |
Kimi K2.6 Moonshot AI · Apr 2026 |
GPT-5.6 Sol OpenAI |
Claude Opus 4.8 Anthropic |
Claude Fable 5 Anthropic |
GLM-5.2 Z.AI |
|
|---|---|---|---|---|---|---|---|
| Architecture & Access | |||||||
| Open weights | ✓ Kimi K3 License | ✓ Modified MIT | ✓ Modified MIT | ✗ | ✗ | ✗ | ✓ |
| Parameters (total / active) | 2.8T / 104B | 1T / 32B | 1T / 32B | Undisclosed | Undisclosed | Undisclosed | Undisclosed |
| Context window | 1M (1,048,576) | 256K | 262K | 1M | 200K | 1M | - |
| Multimodal (image/video) | ✓ + native video | ✓ MoonViT | ✓ | ✓ | ✓ | ✓ | ✓ |
| Benchmarks (vendor model cards · Sep 2026) | |||||||
| Terminal-Bench 2.1 | 88.3 | - | - | 88.8 | 85.0 | 88.0 | 82.7 |
| FrontierSWE | 81.2 | - | - | 71.3 | - | 86.6 | 67.3 |
| Toolathlon-Verified | 76.5 | - | - | 74.9 | - | 77.9 | 59.9 |
| SWE Marathon | 42.0 ★ | - | - | - | 40.0 | 35.0 | - |
| GPQA Diamond | 93.5% ★ | - | - | ~93% | ~91% | ~94% | - |
| LMArena Frontend Code | #1 · 1,679 Elo ★ | - | - | Top 5 | Top 5 | Top 3 | - |
| GDPval-AA v2 (Elo) | 1,686 | - | - | 1,736 | - | 1,747 | 1,510 |
| SWE-bench Verified | Pending (indep.) | Pending | 80.2% ✓ | ~60% | 88.6% | 95.0% | - |
| MCP Mark Verified | - | 81.1 ★ | ~73 | - | 76.4 | - | - |
| Artificial Analysis Index v4.3 (Sep 7) | 44 | - | - | - | - | - | 45 (GLM-5.3) |
| Speed & Cost | |||||||
| API input $/1M | $3.00 | $0.95 | $0.55 | ~$5.00 | $15.00 | ~$20.00 | - |
| API output $/1M | $15.00 | $4.00 | $2.65 | ~$60.00 | $25.00 | $50.00 | - |
| HighSpeed serving | Always thinking | 260 tok/s ⚡ (HS) | ~60 tok/s | ~80 tok/s | ~60 tok/s | ~50 tok/s | - |
| Platform & Ecosystem | |||||||
| Agent Swarm (parallel agents) | Swarm Max | via K2.6 | 300 agents | Limited | Limited | Limited | Limited |
| In GitHub Copilot picker | ✓ Since Aug 6 (Fireworks) | ✓ Since Jul 1 (Azure) | - | ✓ | ✓ | ✓ | - |
| Integrated office tools | Full suite | Docs/Slides/Sheets/Code | Full suite | Canvas | Artifacts | Artifacts | - |
| Free tier with agent mode | ✓ 3/day | ✓ | ✓ | Limited | Limited | Limited | Limited |
★ = vendor-published at launch. Terminal-Bench, FrontierSWE, Toolathlon and GDPval figures from the Kimi-K3 model card (updated Sep 1, 2026) and competitor cards; SWE-bench Verified for Opus 4.8 / Fable 5 from Anthropic. Throughput and competitor pricing approximate. Always verify current benchmarks and pricing on official pages.