Models/Collections/Automate visual understanding
Blueprint4 ModelsUpdated Dec 2025

Automate visual understanding

Captioning, visual QA, video understanding, and segmentation for analysis and moderation workflows.

Vision stack.Describe, answer, segment

DescribeCaptioningVisual QA

Qwen2.5-VL-7B-Instruct

Instruction-tuned multimodal vision-language model

$0.0019/95 - 105 tokens
Run

Qwen2.5-VL-7B-Instruct handles captioning, visual reasoning, question answering, and structured output generation over images.

IMG → TXTCAPTION

Tradeoff. Token-based pricing lands around $0.0019 for typical caption lengths. A 7B backbone; escalate hard cases to a frontier model.

Open weights.Hosted alternatives for this stack

Open-weight models covering the same tasks as this stack, running on Runware's own optimized compute and billed on compute time. License terms vary per model — check each model page before self-hosting.

Google

Gemma 4 31B

by Google

Open 31B multimodal reasoning model for coding, long context, and agentic workflows

TXT → TXTIMG → TXT
from $0.102/Input tokens / 1M
Moonshot AI

Kimi K2.6

by Moonshot AI

Open frontier multimodal LLM for coding, long-horizon execution, and tool-rich workflows

TXT → TXTIMG → TXT
from $0.6/M/Input tokens / 1M
DeepSeek

DeepSeek-V4-Flash-0731

by DeepSeek

Fast frontier LLM with 1M context, tool use, and dual thinking modes

TXT → TXT
from $0.076/Input tokens / 1M
DeepSeek

DeepSeek-V4-Pro

by DeepSeek

High-capability frontier LLM with 1M context, stronger agent performance, and dual thinking modes

TXT → TXT
from $0.961/Input tokens / 1M

Common questions

Describe — Qwen2.5-VL-7B-Instruct; Analyze — Gemini 3 Flash; Transcribe — Memories Video Captioning; Segment — YOLOv8s Person Seg. Each row above expands with the model's live pricing, capability chips, and sample outputs.

Qwen2.5-VL-7B-Instruct and YOLOv8s Person Seg run on Runware's own optimized compute (the platform's open-weight tier, billed on compute time) — check each model page for weight availability and license terms before self-hosting. The remaining picks are partner-served.

The lowest listed starting price in this stack is Qwen2.5-VL-7B-Instruct from $0.0019 per 95 - 105 tokens; prices are read live from the catalog and scale with resolution, duration, and quality tier. The Production notes below cover how pricing is measured.

Tradeoffs to consider

Memories Video Captioning and YOLOv8s Person Seg ship no per-call price in the catalog today. Verify their cost on the model pages before sizing batch workloads; the captioning and multimodal stages have published token-based rates.

Qwen2.5-VL at roughly $0.0019 per caption handles the bulk of description and QA work; Gemini 3 Flash at $0.5 per 1M input tokens adds video and audio input plus harder reasoning. Route by difficulty rather than defaulting everything to the frontier tier.

Production notes

Every price on this page is the model's published rate from the live Runware catalog, using the cheapest listed configuration unless stated otherwise. Prices vary with resolution, duration, quality tier, or token volume, so check the pricing table on each model page before estimating unit economics.

The Runware catalog does not publish per-model latency figures, so this page does not quote end-to-end timings. Where a model's own description commits to speed (for example sub-second generation or realtime streaming), that claim is repeated here. For anything else, benchmark the exact models in the Playground with your own payload sizes before committing to an SLA.

Models are addressed by versioned AIR identifiers, so a workflow pinned to specific model versions keeps producing the same behaviour as new versions ship. Adopt upgrades deliberately by re-running your evaluation set against the new version before switching production traffic.

About this collection

A blueprint bundles the small set of models you would actually wire together for one production use case, instead of ranking a whole category. Each pick covers one stage of the workflow and links to its model page for full schema, pricing, and examples.

Picks are live models from the Best Captioning (5 models) and Best Masking (15) collections plus the catalog's video-to-text capability group (9 live models ship io:video-to-text), covering image captioning, multimodal understanding, video transcription, and segmentation masks.