Best LLMs
Large language models for general-purpose text generation, reasoning, summarization, and conversation. Capable of understanding complex instructions and producing high-quality, coherent text responses.
Best rated
by Moonshot AI
Kimi K3 is Moonshot AI's flagship model for frontier reasoning, software engineering, and knowledge-intensive work. It combines a 1M-token context window with native visual understanding, configurable reasoning effort, structured outputs, tool calling, and strong long-horizon execution across large codebases and multi-step research workflows. It is well suited to demanding coding agents, deep research systems, document-heavy assistants, and other production applications that need sustained reasoning over very large contexts.
Featured Models
Top-performing models in this category, recommended by our community and performance benchmarks.
by Anthropic
Claude Fable 5 is Anthropic's new top-capability generally available Claude model. It is built for long-running coding, agentic execution, multimodal reasoning, research, and high-stakes professional workflows, with stronger long-horizon performance than earlier Opus models, state-of-the-art vision, and a 1M-token context window. Safety classifiers can route some sensitive requests to Claude Opus 4.8 instead.
by DeepSeek
DeepSeek-V4-Pro is DeepSeek's flagship V4 language model for coding, reasoning, and agent workflows that need stronger overall capability than the Flash variant. It supports both thinking and non-thinking modes, a 1M token context window, up to 384K output tokens, tool calls, JSON output, and long-context operation for demanding software, research, and structured professional workloads.
by MiniMax
MiniMax M3 is MiniMax's new flagship open-weight model for coding, agentic execution, and multimodal reasoning. It is built around the MiniMax Sparse Attention architecture, supports up to 1 million tokens of context, accepts image and video input in addition to text, and is designed for long-running software, research, browsing, and desktop-operation workflows that need strong tool use and sustained multi-step performance.
by Google
Gemini 3.6 Flash is an updated multimodal reasoning model in Google's Flash series for production agent workflows, coding, and knowledge-intensive tasks. It is designed to deliver stronger coding quality, better computer-use performance, and more efficient reasoning than the previous generation, with fewer unnecessary steps and lower token usage while preserving the Flash family's balance of speed, scale, and capability.
by Google
Gemini 3.5 Flash is Google’s most intelligent Flash-series multimodal model for sustained frontier performance on agentic and coding tasks. It accepts text, images, video, audio, and PDFs, and is designed for long-horizon workflows, sub-agent orchestration, complex coding loops, multimodal understanding, and high-speed reasoning at production scale.
by Google
Gemini 3.5 Flash-Lite is a lightweight multimodal model in Google's Gemini 3.5 family, built for high-throughput and latency-sensitive production workloads. It is well suited to agentic search, document processing, coding subtasks, routing, summarization, and other flows where fast response times and scale matter as much as model quality. It combines the Lite tier's efficiency profile with stronger long-context, coding, and agent-style performance than earlier Flash-Lite generations.
by Google
Gemini 3.5 Flash Cyber is a cybersecurity-specialized model built on top of Gemini 3.5 Flash and tuned for software security workflows. It is intended for defensive use cases such as vulnerability discovery, validation, patch generation, and remediation support, with an emphasis on efficient security analysis rather than broad general-purpose assistant tasks.
by Anthropic
Claude Opus 4.8 is Anthropic's highest-capability Claude model. It is built for demanding coding, agent orchestration, multimodal reasoning, and professional workflows that need strong instruction following, adaptive and extended thinking, high-resolution vision, and a 1M-token context window.
by xAI
Grok 4.3 is xAI's flagship language model for agentic reasoning, strong instruction following, and minimal hallucinations. It supports text and image input, a 1 million token context window, configurable reasoning effort including non-reasoning mode, function calling, and structured outputs for production assistants, coding workflows, and long-context analysis.
by DeepSeek
DeepSeek-V4-Flash is DeepSeek's fast, efficient, and cost-focused frontier language model for coding, reasoning, and agent workflows. It supports both thinking and non-thinking modes, a 1M token context window, up to 384K output tokens, tool calls, JSON output, and efficient long-context operation for software, research, and structured professional tasks.
by Google
Gemma 4 31B is Google's flagship dense open-weights model in the Gemma 4 family. It combines strong reasoning, coding performance, native function calling, multimodal understanding across text, image, and video, and a 256K context window in a 31B-parameter open model designed for local and cloud deployment.
by Moonshot AI
Kimi K2.6 is Moonshot AI's latest flagship open model for coding, reasoning, multimodal understanding, and agentic execution. It is designed for long-horizon software tasks, reliable tool use, autonomous multi-step workflows, coordinated agent swarms, and visual understanding across image and video inputs in addition to text.
by OpenAI
GPT-5.6 Sol is the flagship model in OpenAI's GPT-5.6 family for complex professional work. It combines the strongest reasoning tier in this family with a 1,050,000 token context window, up to 128,000 output tokens, image input, structured outputs, function calling, and broad tool use for demanding coding, research, analysis, and agent workflows.
by OpenAI
GPT-5.6 Terra is the balanced model in OpenAI's GPT-5.6 family, designed for production workloads that need strong reasoning without the cost of the flagship tier. It supports a 1,050,000 token context window, up to 128,000 output tokens, image input, structured outputs, function calling, and broad tool use, making it well suited to coding, analysis, extraction, workflow automation, and general-purpose professional assistants.
by OpenAI
GPT-5.6 Luna is the most cost-efficient model in OpenAI's GPT-5.6 family, aimed at high-volume and cost-sensitive workloads. It supports a 1,050,000 token context window, up to 128,000 output tokens, image input, structured outputs, function calling, and broad tool use for classification, extraction, summarization, routing, tagging, and lightweight agent tasks.
by OpenAI
GPT-5.5 is OpenAI's newest frontier model for complex professional work, with strong performance in coding, reasoning, and tool-using workflows. It supports a 1,050,000 token context window, 128,000 max output tokens, configurable reasoning effort, image input, and a broad tool stack including web search, file search, code interpreter, hosted shell, apply patch, skills, MCP, tool search, and computer use.
by Z.ai
GLM-5.1 is Z.ai’s flagship language model for agentic engineering, coding, reasoning, and tool-driven workflows. It supports a 200K token context window with up to 128K output tokens, deep thinking, function calling, structured output, and streaming tool calls, and is designed to stay effective over long multi-step sessions rather than only short-horizon tasks.
by Google
Gemini 3.1 Pro is Google’s flagship multimodal language model that processes text alongside images, audio, video, code, and documents. It offers high-performance reasoning, complex instruction following, and deep contextual understanding for a wide range of tasks across language, analysis, and problem solving.
by Anthropic
Claude Opus 4.7 is Anthropic's highest-capability generally available Claude model. It is designed for demanding coding, agent orchestration, multimodal reasoning, and high-stakes professional workflows, with stronger instruction following, better high-resolution vision, adaptive thinking, and a 1M-token context window.
by Anthropic
Claude Sonnet 4.6 is Anthropic's most capable Sonnet model, built for daily production use across coding, agent workflows, long-context reasoning, computer use, and professional knowledge work. It supports adaptive and extended thinking, strong instruction following, high-volume automation, and a 1M-token context window in beta.
by MiniMax
MiniMax M2.7 is a long-context LLM designed for agentic workflows across software engineering, search and tool use, and high-value office productivity tasks. It’s built for multi-step execution, with strong instruction following and dependable task decomposition, making it a solid default for production assistants that write code, call tools, and handle complex document workflows.
by MiniMax
MiniMax M2.7-Highspeed is the performance-tuned variant of M2.7, built for lower latency and higher throughput while keeping output behavior consistent with the standard model. It’s a strong fit for interactive coding agents, tool-calling pipelines, and office automation flows where responsiveness matters.
by Anthropic
Claude Haiku 4.5 is Anthropic's fastest and most cost-efficient Claude model. It is built for latency-sensitive applications, high-volume agents, sub-agent orchestration, coding assistance, and budget-conscious deployments that still need strong reasoning and multimodal understanding.
by Google
Gemini 3 Flash is Google’s flagship multimodal language model that processes text alongside images, audio, video, code, and documents. It offers high-performance reasoning, complex instruction following, and deep contextual understanding for a wide range of tasks across language, analysis, and problem solving.
by Google
Gemini 3.1 Flash Lite is Google’s flagship multimodal language model that processes text alongside images, audio, video, code, and documents. It offers high-performance reasoning, complex instruction following, and deep contextual understanding for a wide range of tasks across language, analysis, and problem solving
by OpenAI
GPT-5.4 is OpenAI's flagship large language model, featuring a 1 million token context window, native computer use, and a 33% reduction in factual errors over GPT-5.2. It integrates coding capabilities from GPT-5.3-Codex, is 47% more token-efficient, and supports configurable reasoning effort for complex professional tasks.
by OpenAI
GPT-5.4 Pro is the high-performance variant of GPT-5.4, optimized for enterprise-grade professional tasks. It offers deeper reasoning, enhanced accuracy, and extended compute for complex multi-step workflows including document creation, spreadsheet analysis, and autonomous agent orchestration. It shares the 1 million token context window and native computer use capabilities of the standard GPT-5.4.
by OpenAI
GPT-5.4 Mini is a compact, efficient variant of GPT-5.4 designed for coding assistants, subagent orchestration, and multimodal applications requiring faster responsiveness. It supports a 400K token context window and retains native computer use and configurable reasoning effort at a lower cost than the flagship model.
by OpenAI
GPT-5.4 Nano is the smallest and fastest variant of GPT-5.4, designed for high-throughput, low-latency tasks such as classification, data extraction, ranking, and lightweight automation. It prioritizes speed and cost efficiency for simple, high-volume workloads and is available exclusively via the API.





























