Kimi K3

Coming soon to Runware

Moonshot's 2.8-trillion-parameter frontier LLM, available on Runware for up to 50% less than other providers

1M-token context and native image and video input, with reasoning that never switches off. Register below to run it here first.

2.8TParameters
1MToken context
88.3Terminal-Bench 2.1
-50%vs other providers

Every workflow, one endpoint.

Chat & reasoningAgentic codingVision understanding1M-token contextTool use

One API, OpenAI-compatible endpoint. Point your existing SDK at Runware and go.

Kimi K3 Benchmarks

Moonshot's published evals, thinking effort maxed for every model. Kimi K3 leads outright on agentic coding endurance and browsing, and effectively ties GPT-5.6 Sol on terminal work, at $3 per million input tokens and $0.30 on cache hits.

BenchmarkKimi K3GPT-5.6 SolFable 5Opus 4.8GPT-5.5GLM-5.2
Terminal-Bench 2.188.388.884.684.683.482.7
Program Bench77.877.676.871.970.863.7
SWE Marathon42.039.035.040.014.013.0
BrowseComp91.290.488.084.384.4

Moonshot-published figures, all models at max thinking effort. Fable 5 results include potential fallbacks; GPT-5.6 Sol results include potential cyberguards. Dash: not reported.

Built for real workloads

Long-running agents and oversized codebases, plus pipelines that previously required separate vision and reranking models.

Coding agents

Multi-step engineering across large codebases; K3 leads Program Bench and SWE Marathon and beats Opus 4.8 on Terminal-Bench 2.1.

Long-document analysis

Entire contract libraries and research archives: 1M tokens fits them in one pass, no chunking, no retrieval layer.

Multimodal reasoning

Visual debugging and QA flows, with screenshots and video reasoned over natively.

Autonomous agents

Loads new tools mid-conversation via system messages and enforces strict JSON Schema output, so long-horizon loops stay on rails.

High-volume inference

Automatic context caching with no configuration needed: hit rates above 90% on coding workloads and input cost drops to $0.30/M on hits.

Drop-in migration

OpenAI-compatible: change the base URL and model ID; streaming and tool calls carry over untouched, as does structured output.

Kimi K3 pricing

Runware will serve Kimi K3 at up to 50% less than any other inference provider. Exact prices publish at launch; Moonshot's list prices below are the ceiling, and automatic context caching drops input cost to $0.30/M on hits.

$3.00/M tokensInput (Moonshot list)
$0.30/M tokensInput on cache hits
$15.00/M tokensOutput (Moonshot list)

Register for early access and we'll send you Runware's exact Kimi K3 prices before they go public.

Get launch pricing first

Frequently asked questions

What is Kimi K3?

Moonshot AI's frontier LLM, released July 16, 2026: 2.8 trillion parameters with 16 of 896 experts active per pass, a 1M-token context window and native multimodal input, with reasoning always on. Weights go open July 27, the largest open-weight release ever.

How much will it cost on Runware?

Up to 50% less than any other inference provider: Moonshot lists $3 per million input tokens ($0.30 on cache hits) and $15 per million output; Runware's exact prices publish at launch. Register below and we'll send them to you first.

What makes K3 different?

Open weights and a 1M-token context window, combined with frontier agentic coding performance: 88.3 on Terminal-Bench 2.1, above Claude Opus 4.8's 84.6. Thinking mode can't be switched off, only tuned, and automatic context caching hits over 90% on coding workloads.

When will it be live on Runware?

We're integrating it now. Register below and you'll be notified at launch, ahead of public availability.

Do I need a new integration?

No. The endpoint is OpenAI-compatible: change the base URL and model ID, and streaming, tool calls, and structured output carry over unchanged.

What workloads is it best for?

Agentic coding, long-document analysis, multimodal reasoning, and long-horizon agents: it loads new tools mid-conversation and enforces strict JSON Schema output, so autonomous loops stay on rails.

Migration offer

Bring your workload. The switch is on us.

Already running an LLM in production? We'll credit your Runware account so you can replay your real traffic on Kimi K3 and compare cost and quality against your current provider, before you spend a cent of your own.

  1. 1Tell us your current provider and monthly volume
  2. 2We set you up with testing credits
  3. 3Run your real prompts, keep what wins
Claim migration credits

Claim your migration credits

Tell us your current provider and use case along with your monthly volume. We'll set you up with testing credits and dedicated RPM so you can benchmark Kimi K3 on Runware against what you run today.