Z.ai
Z.ai

GLM-5.3-Flash

Efficient multimodal GLM for coding, agents, and long-context visual understanding

Available via OpenAI-compatible API

GLM-5.3-Flash Overview

GLM-5.3-Flash is Z.ai's lower-cost multimodal model for coding, long-horizon agent workflows, and visual reasoning. It is the first natively multimodal model in the GLM-5 series, combining text, image, and video understanding with stronger coding and agent performance than earlier GLM flash-tier behavior would suggest. Z.ai positions it as a default model for frequent professional work that needs fast feedback, efficient long-context serving, and reliable reasoning across software, documents, charts, interfaces, screenshots, and other visually grounded tasks.

Token based
Input tokens / 1M$0.15
Output tokens / 1M$0.5
Cache read$0.03

Commercial use

More models from Z.ai

GLM-5.3

Coming Soon

GLM-5.3 is Z.ai's flagship language model for coding, tool-using agents, and sustained multi-step problem solving. It builds on the long-horizon engineering focus of the recent GLM line with stronger software engineering performance, better benchmark results across terminal and agent evaluations, and notably stronger capability in security-oriented code analysis and vulnerability discovery. It is a strong fit for demanding coding assistants, automation systems, and engineering workflows that need reliable reasoning over large task context.

GLM-5.2

Openai Compatible

GLM-5.2 is Z.ai's flagship language model for long-horizon coding, agentic engineering, and sustained multi-step execution. It is designed to keep large project context coherent over extended runs, with a 1M token context window, 128K max output, multiple thinking modes, function calling, structured output, context caching, streaming, and MCP support for tool-rich workflows.

GLM-5.1 is Z.ai’s flagship language model for agentic engineering, coding, reasoning, and tool-driven workflows. It supports a 200K token context window with up to 128K output tokens, deep thinking, function calling, structured output, and streaming tool calls, and is designed to stay effective over long multi-step sessions rather than only short-horizon tasks.