
DeepSeek-V4.1-Flash
Multimodal 1M-context MoE for efficient coding, reasoning, and long-horizon agent workflows
DeepSeek-V4.1-Flash
Multimodal 1M-context MoE for efficient coding, reasoning, and long-horizon agent workflows
DeepSeek-V4.1-Flash Overview
DeepSeek-V4.1-Flash is a multimodal mixture-of-experts model for coding, reasoning, visual understanding, and tool-using agents. It processes text and images, generates text, and combines a 1M-token context window with continuously adjustable reasoning effort. Its Causal Encoder-Decoder architecture, sparse attention, and compressed KV cache reduce the memory and computation required for long inputs, making it well suited to large-codebase work, document analysis, research, and sustained agentic tasks.
More models from DeepSeek
DeepSeek-V4-Flash is DeepSeek's fast, efficient, and cost-focused frontier language model for coding, reasoning, and agent workflows. It supports both thinking and non-thinking modes, a 1M token context window, up to 384K output tokens, tool calls, JSON output, and efficient long-context operation for software, research, and structured professional tasks.
DeepSeek-V4-Pro is DeepSeek's flagship V4 language model for coding, reasoning, and agent workflows that need stronger overall capability than the Flash variant. It supports both thinking and non-thinking modes, a 1M token context window, up to 384K output tokens, tool calls, JSON output, and long-context operation for demanding software, research, and structured professional workloads.

