Kimi K3
Coming soon to RunwareMoonshot's 2.8-trillion-parameter frontier LLM, available on Runware for up to 50% less than other providers
1M-token context and native image and video input, with reasoning that never switches off. Register below to run it here first.
Every workflow, one endpoint.
One API, OpenAI-compatible endpoint. Point your existing SDK at Runware and go.
Kimi K3 Benchmarks
Moonshot's published evals, thinking effort maxed for every model. Kimi K3 leads outright on agentic coding endurance and browsing, and effectively ties GPT-5.6 Sol on terminal work, at $3 per million input tokens and $0.30 on cache hits.
| Benchmark | Kimi K3 | GPT-5.6 Sol | Fable 5 | Opus 4.8 | GPT-5.5 | GLM-5.2 |
|---|---|---|---|---|---|---|
| Terminal-Bench 2.1 | 88.3 | 88.8 | 84.6 | 84.6 | 83.4 | 82.7 |
| Program Bench | 77.8 | 77.6 | 76.8 | 71.9 | 70.8 | 63.7 |
| SWE Marathon | 42.0 | 39.0 | 35.0 | 40.0 | 14.0 | 13.0 |
| BrowseComp | 91.2 | 90.4 | 88.0 | 84.3 | 84.4 | — |
Moonshot-published figures, all models at max thinking effort. Fable 5 results include potential fallbacks; GPT-5.6 Sol results include potential cyberguards. Dash: not reported.
Built for real workloads
Long-running agents and oversized codebases, plus pipelines that previously required separate vision and reranking models.
Coding agents
Multi-step engineering across large codebases; K3 leads Program Bench and SWE Marathon and beats Opus 4.8 on Terminal-Bench 2.1.
Long-document analysis
Entire contract libraries and research archives: 1M tokens fits them in one pass, no chunking, no retrieval layer.
Multimodal reasoning
Visual debugging and QA flows, with screenshots and video reasoned over natively.
Autonomous agents
Loads new tools mid-conversation via system messages and enforces strict JSON Schema output, so long-horizon loops stay on rails.
High-volume inference
Automatic context caching with no configuration needed: hit rates above 90% on coding workloads and input cost drops to $0.30/M on hits.
Drop-in migration
OpenAI-compatible: change the base URL and model ID; streaming and tool calls carry over untouched, as does structured output.
Kimi K3 pricing
Runware will serve Kimi K3 at up to 50% less than any other inference provider. Exact prices publish at launch; Moonshot's list prices below are the ceiling, and automatic context caching drops input cost to $0.30/M on hits.
Register for early access and we'll send you Runware's exact Kimi K3 prices before they go public.
Get launch pricing firstFrequently asked questions
What is Kimi K3?
Moonshot AI's frontier LLM, released July 16, 2026: 2.8 trillion parameters with 16 of 896 experts active per pass, a 1M-token context window and native multimodal input, with reasoning always on. Weights go open July 27, the largest open-weight release ever.
How much will it cost on Runware?
Up to 50% less than any other inference provider: Moonshot lists $3 per million input tokens ($0.30 on cache hits) and $15 per million output; Runware's exact prices publish at launch. Register below and we'll send them to you first.
What makes K3 different?
Open weights and a 1M-token context window, combined with frontier agentic coding performance: 88.3 on Terminal-Bench 2.1, above Claude Opus 4.8's 84.6. Thinking mode can't be switched off, only tuned, and automatic context caching hits over 90% on coding workloads.
When will it be live on Runware?
We're integrating it now. Register below and you'll be notified at launch, ahead of public availability.
Do I need a new integration?
No. The endpoint is OpenAI-compatible: change the base URL and model ID, and streaming, tool calls, and structured output carry over unchanged.
What workloads is it best for?
Agentic coding, long-document analysis, multimodal reasoning, and long-horizon agents: it loads new tools mid-conversation and enforces strict JSON Schema output, so autonomous loops stay on rails.
Migration offer
Bring your workload. The switch is on us.
Already running an LLM in production? We'll credit your Runware account so you can replay your real traffic on Kimi K3 and compare cost and quality against your current provider, before you spend a cent of your own.
- 1Tell us your current provider and monthly volume
- 2We set you up with testing credits
- 3Run your real prompts, keep what wins
Claim your migration credits
Tell us your current provider and use case along with your monthly volume. We'll set you up with testing credits and dedicated RPM so you can benchmark Kimi K3 on Runware against what you run today.