Best Coding Agents
Models optimized for coding, agentic tool use, and multi-step workflows. Ideal for writing, debugging, and refactoring code, executing tool calls, and orchestrating complex long-running automated pipelines.
Best rated
by Moonshot AI
Kimi K3 is Moonshot AI's flagship model for frontier reasoning, software engineering, and knowledge-intensive work. It combines a 1M-token context window with native visual understanding, configurable reasoning effort, structured outputs, tool calling, and strong long-horizon execution across large codebases and multi-step research workflows. It is well suited to demanding coding agents, deep research systems, document-heavy assistants, and other production applications that need sustained reasoning over very large contexts.
Featured Models
Top-performing models in this category, recommended by our community and performance benchmarks.
by Google
Gemini 3.6 Flash is an updated multimodal reasoning model in Google's Flash series for production agent workflows, coding, and knowledge-intensive tasks. It is designed to deliver stronger coding quality, better computer-use performance, and more efficient reasoning than the previous generation, with fewer unnecessary steps and lower token usage while preserving the Flash family's balance of speed, scale, and capability.
by Google
Gemini 3.5 Flash Cyber is a cybersecurity-specialized model built on top of Gemini 3.5 Flash and tuned for software security workflows. It is intended for defensive use cases such as vulnerability discovery, validation, patch generation, and remediation support, with an emphasis on efficient security analysis rather than broad general-purpose assistant tasks.
by OpenAI
GPT-5.6 Sol is the flagship model in OpenAI's GPT-5.6 family for complex professional work. It combines the strongest reasoning tier in this family with a 1,050,000 token context window, up to 128,000 output tokens, image input, structured outputs, function calling, and broad tool use for demanding coding, research, analysis, and agent workflows.
by OpenAI
GPT-5.6 Terra is the balanced model in OpenAI's GPT-5.6 family, designed for production workloads that need strong reasoning without the cost of the flagship tier. It supports a 1,050,000 token context window, up to 128,000 output tokens, image input, structured outputs, function calling, and broad tool use, making it well suited to coding, analysis, extraction, workflow automation, and general-purpose professional assistants.
by MiniMax
MiniMax M2.7-Highspeed is the performance-tuned variant of M2.7, built for lower latency and higher throughput while keeping output behavior consistent with the standard model. It’s a strong fit for interactive coding agents, tool-calling pipelines, and office automation flows where responsiveness matters.
by Google
Gemini 3.5 Flash-Lite is a lightweight multimodal model in Google's Gemini 3.5 family, built for high-throughput and latency-sensitive production workloads. It is well suited to agentic search, document processing, coding subtasks, routing, summarization, and other flows where fast response times and scale matter as much as model quality. It combines the Lite tier's efficiency profile with stronger long-context, coding, and agent-style performance than earlier Flash-Lite generations.
by Google
Gemini 3.1 Flash Lite is Google’s flagship multimodal language model that processes text alongside images, audio, video, code, and documents. It offers high-performance reasoning, complex instruction following, and deep contextual understanding for a wide range of tasks across language, analysis, and problem solving
by Google
Gemini 3 Flash is Google’s flagship multimodal language model that processes text alongside images, audio, video, code, and documents. It offers high-performance reasoning, complex instruction following, and deep contextual understanding for a wide range of tasks across language, analysis, and problem solving.
by MiniMax
MiniMax M2.7 is a long-context LLM designed for agentic workflows across software engineering, search and tool use, and high-value office productivity tasks. It’s built for multi-step execution, with strong instruction following and dependable task decomposition, making it a solid default for production assistants that write code, call tools, and handle complex document workflows.
by Google
Gemini 3.1 Pro is Google’s flagship multimodal language model that processes text alongside images, audio, video, code, and documents. It offers high-performance reasoning, complex instruction following, and deep contextual understanding for a wide range of tasks across language, analysis, and problem solving.
by OpenAI
GPT-5.4 Pro is the high-performance variant of GPT-5.4, optimized for enterprise-grade professional tasks. It offers deeper reasoning, enhanced accuracy, and extended compute for complex multi-step workflows including document creation, spreadsheet analysis, and autonomous agent orchestration. It shares the 1 million token context window and native computer use capabilities of the standard GPT-5.4.
by OpenAI
GPT-5.4 is OpenAI's flagship large language model, featuring a 1 million token context window, native computer use, and a 33% reduction in factual errors over GPT-5.2. It integrates coding capabilities from GPT-5.3-Codex, is 47% more token-efficient, and supports configurable reasoning effort for complex professional tasks.
by OpenAI
GPT-5.4 Mini is a compact, efficient variant of GPT-5.4 designed for coding assistants, subagent orchestration, and multimodal applications requiring faster responsiveness. It supports a 400K token context window and retains native computer use and configurable reasoning effort at a lower cost than the flagship model.
by OpenAI
GPT-5.4 Nano is the smallest and fastest variant of GPT-5.4, designed for high-throughput, low-latency tasks such as classification, data extraction, ranking, and lightweight automation. It prioritizes speed and cost efficiency for simple, high-volume workloads and is available exclusively via the API.














