Models/Collections/Best Coding Agents
Capability · Text (LLMs)19 ModelsUpdated Sep 2026

Best Coding Agents

Models optimized for coding, agentic tool use, and multi-step workflows. Ideal for writing, debugging, and refactoring code, executing tool calls, and orchestrating complex long-running automated pipelines.

The picks.Reliable models for coding, tool use, and agentic workflows

TXT → TXTIMG → TXT

Kimi K3

Flagship multimodal reasoning LLM for long-horizon coding, knowledge work, and tool-rich agent workflows

compatible APITXT → TXTIMG → TXT

Open weights.Hosted alternatives for this stack

Open-weight models covering the same tasks as this stack, running on Runware's own optimized compute and billed on compute time. License terms vary per model — check each model page before self-hosting.

Google

Gemma 4 31B

by Google

Open 31B multimodal reasoning model for coding, long context, and agentic workflows

TXT → TXTIMG → TXT
from $0.000567/text-generation
M

Shieldstral 1.0 3B

by mistral ai

Compact multimodal moderation model that judges text and images against your own policy

TXT → TXTIMG → TXT
Compute-time pricing
L

Laya

by Independent

Open-weight System 1 model for low-latency typed decisions without free-form generation

TXT → TXTCLASSIFICATION
Compute-time pricing
Z.ai

GLM-5.3

by Z.ai

Flagship coding and agentic LLM with stronger long-horizon execution and security-oriented reasoning

TXT → TXT
Compute-time pricing

Common questions

Yes — Kimi K3, GLM-5.3-Flash, GLM-5.2, DeepSeek-V4-Pro-0423 and 2 more run on Runware's own optimized compute (the platform's open-weight tier, billed on compute time). Check each model page for license terms before self-hosting.

GPT-6 Sol, released September 2026 per the live catalog. Membership updates automatically as the catalog publishes new models to this collection.

Production notes

Every price on this page is the model's published rate from the live Runware catalog, using the cheapest listed configuration unless stated otherwise. Prices vary with resolution, duration, quality tier, or token volume, so check the pricing table on each model page before estimating unit economics.

The Runware catalog does not publish per-model latency figures, so this page does not quote end-to-end timings. Where a model's own description commits to speed (for example sub-second generation or realtime streaming), that claim is repeated here. For anything else, benchmark the exact models in the Playground with your own payload sizes before committing to an SLA.

Models are addressed by versioned AIR identifiers, so a workflow pinned to specific model versions keeps producing the same behaviour as new versions ship. Adopt upgrades deliberately by re-running your evaluation set against the new version before switching production traffic.

About this collection

Models optimized for coding, agentic tool use, and multi-step workflows. Ideal for writing, debugging, and refactoring code, executing tool calls, and orchestrating complex long-running automated pipelines.

Membership comes directly from the Runware catalog: the 19 models on this page are the live catalog's own membership for the "Best Coding Agents" collection. Names, descriptions, pricing, capability chips, samples, and guides are all read live from the catalog — nothing here is hand-curated.