Kimi K3

Flagship multimodal reasoning LLM for long-horizon coding, knowledge work, and tool-rich agent workflows

Kimi K3 Overview

Kimi K3 is Moonshot AI's flagship model for frontier reasoning, software engineering, and knowledge-intensive work. It combines a 1M-token context window with native visual understanding, configurable reasoning effort, structured outputs, tool calling, and strong long-horizon execution across large codebases and multi-step research workflows. It is well suited to demanding coding agents, deep research systems, document-heavy assistants, and other production applications that need sustained reasoning over very large contexts.

Token based
Input tokens / 1M$3
Output tokens / 1M$15
Cache Read / 1M$0.3

Commercial use

How to Use Kimi K3

Overview

Kimi K3 is Moonshot AI's flagship multimodal reasoning model for long-horizon coding, knowledge work, and tool-using agent systems.

It combines a 1M-token context window with native visual understanding, structured outputs, configurable reasoning effort, and strong support for multi-step execution across large codebases and document-heavy workflows.

Strengths

Long-Horizon Coding

Kimi K3 is built for software tasks that unfold over many steps rather than a single exchange. It is a strong fit for debugging, refactoring, architecture changes, repository-wide edits, and other engineering workflows that require sustained reasoning across large codebases.

Knowledge-Intensive Work

The model is designed for research, analysis, and other professional workflows where large amounts of context need to stay in scope. It is well suited to document-heavy assistants, strategic analysis, synthesis over many sources, and multi-step knowledge work.

Multimodal Understanding

Kimi K3 supports image and video input in addition to text. That makes it useful for workflows involving screenshots, UI states, diagrams, visual documents, and video content that must be analyzed as part of a broader reasoning task.

Strong Tool Use

K3 supports tool calling and is positioned for agent-style workflows that combine reasoning with external actions. It works well in systems that need planning, execution, tool orchestration, and recovery over multiple turns.

Structured Outputs

The model supports strict JSON-schema-constrained outputs, which makes it easier to use in production systems that depend on reliable machine-readable responses.

Capabilities

Text-to-Text

Kimi K3 handles general language tasks including coding assistance, summarization, planning, reasoning, drafting, transformation, and research-oriented outputs.

Image-to-Text

The model can understand image input and use it as part of broader reasoning and analysis workflows.

Video-to-Text

The model can take video input for analysis and description tasks, making it useful for workflows that depend on temporal visual content rather than only still images.

Tool Calling

K3 is a strong fit for function calling, agent orchestration, and external-action pipelines that combine reasoning with tools.

Long-Context Work

Kimi K3 supports a 1,000,000-token context window, making it suitable for workflows where very large context must remain in scope across a long task.

Input and Output

  • AIR ID: moonshotai:kimi@k3
  • Input: text, images, and video
  • Output: text
  • Context window: 1,000,000 tokens
  • Reasoning effort: supported
  • Tool use: supported
  • Structured output: supported
  • Partial mode: supported

Best Fit

  • Coding agents
  • Long-horizon engineering tasks
  • Tool-using assistants
  • Deep research and analysis workflows
  • Multimodal reasoning over images and video
  • Structured professional automation