FLUX 3 Video

Live now on Runware

Black Forest Labs' first video generation model

Access FLUX 3 Video, Black Forest Labs' first video generation model, trained jointly across image, video, audio, and action for physically coherent scenes with sound generated alongside picture.

20sMax clip length
1080pMax resolution
7Aspect ratios
3Generation modes

Three ways to build a video

FLUX 3 Video can generate an entirely new clip from a written prompt, animate forward from a single image, or follow keyframes pinned to specific points in the sequence.

Text to video

A prompt alone is enough. FLUX 3 Video generates an entirely new clip from scratch, with audio synchronized to the action as it happens.

Image to video

Start from a single image and animate forward from it. The source frame keeps its look while the prompt drives the motion and camera work, with sound layered in automatically.

Keyframes

Pin the frames that matter, such as where a shot should begin or end, and FLUX 3 Video fills in the motion between them. It plays more like storyboarding than prompting.

Why FLUX 3 Video stands out

One model, not a pipeline

Most video systems stitch together separate models for picture and sound. FLUX 3 Video was trained jointly across image, video, audio, and action, so it learns how they relate instead of guessing after the fact.

Audio generated by default

Synchronized audio, including native multilingual dialogue, comes on by default and can be switched off per request when it isn't needed.

Full stylistic range in one model

Candid camcorder footage, animation, motion design, and cinematic photoreal all come from the same model, so switching styles doesn't mean switching tools.

Character consistency across cuts

When a generation includes multiple shots, the same character or subject holds up visually from one cut to the next, without extra prompting to keep it consistent.

Multi-shot sequences, clean on-screen text

Generate a full sequence with hard cuts inside one call instead of stitching shots together, and render titles or animated typography cleanly right in the frame.

Built for length and resolution

Generate clips from 5 to 20 seconds at 480p or 1080p, across aspect ratios spanning cinematic 21:9 to vertical 9:21.

Built for real workflows

FLUX 3 Video is Black Forest Labs' first step into video, built from a unified multimodal architecture rather than an image model stretched to fit. It covers the full stylistic range in a single model, from candid camcorder footage to cinematic photoreal, so switching styles doesn't mean switching tools.

Product and brand footage

Turn a single product photo into polished commercial footage, with motion and sound generated to match, and no live shoot required.

Multi-shot cinematic sequences

Generate a full sequence with hard cuts inside one call, so establishing shots and follow-ups come back together with audio that matches the action.

Motion design and on-screen text

Produce stylized motion graphics with clean in-video titles and animated typography, from the same model used for photoreal footage.

Storyboarding with keyframes

Pin the key moments of a shot and let the model generate everything in between, closer to directing than prompting.

Localized dialogue and voiceover

Generate native multilingual dialogue synced to the action, so the same clip can carry a script in more than one language without a separate dubbing pass.

Stylized and handmade looks

Move between candid camcorder footage, animation, and motion design from the same model, so a campaign can shift styles without switching tools.

Frequently asked questions

What is FLUX 3 Video?

Black Forest Labs' multimodal foundation model for video generation with synchronized audio, trained jointly across image, video, audio, and action. It covers text to video and image to video on one architecture, with keyframe control layered on top to pin what happens on screen.

What generation modes does it support?

Text to video and image to video, plus keyframe control to pin specific frames in a sequence. All of it runs through the same Runware integration.

Does it generate audio automatically?

Yes, by default across every mode, and it can be switched off per request. Audio comes from the same joint training that grounds the model's sense of motion, rather than a separate pass added afterward.

What resolutions and durations are supported?

Clips from 5 to 20 seconds at 480p or 1080p, across seven aspect ratios: 21:9, 16:9, 4:3, 1:1, 3:4, 9:16, and 9:21.

Do I need separate integrations for each mode?

No. Text to video and image to video, plus keyframe control, all run through the same Runware API as every other model on the platform, under the same key and billing.

How much does FLUX 3 Video cost?

Pricing is per second of output rather than a flat per-clip rate. Text to video and image to video cost $0.17 per second at 720p and $0.29 per second at 1080p. Keyframe control runs on the same per-second, usage-based pricing.

Is FLUX 3 Video available now?

Yes. FLUX 3 Video is live on Runware, available through the API and Playground under the same key and billing as every other model on the platform.

Talk to us about production volume

Have questions about FLUX 3 Video? Chat to our team about enterprise usage, including volume discounts and dedicated RPM, and we'll follow up shortly.