FLUX 3 Video
Live now on RunwareBlack Forest Labs' first video generation model
Access FLUX 3 Video, Black Forest Labs' first video generation model, trained jointly across image, video, audio, and action for physically coherent scenes with sound generated alongside picture.
Three ways to build a video
FLUX 3 Video can generate an entirely new clip from a written prompt, animate forward from a single image, or follow keyframes pinned to specific points in the sequence.
Text to video
A prompt alone is enough. FLUX 3 Video generates an entirely new clip from scratch, with audio synchronized to the action as it happens.
Image to video
Start from a single image and animate forward from it. The source frame keeps its look while the prompt drives the motion and camera work, with sound layered in automatically.
Keyframes
Pin the frames that matter, such as where a shot should begin or end, and FLUX 3 Video fills in the motion between them. It plays more like storyboarding than prompting.
Why FLUX 3 Video stands out
One model, not a pipeline
Most video systems stitch together separate models for picture and sound. FLUX 3 Video was trained jointly across image, video, audio, and action, so it learns how they relate instead of guessing after the fact.
Audio generated by default
Synchronized audio, including native multilingual dialogue, comes on by default and can be switched off per request when it isn't needed.
Full stylistic range in one model
Candid camcorder footage, animation, motion design, and cinematic photoreal all come from the same model, so switching styles doesn't mean switching tools.
Character consistency across cuts
When a generation includes multiple shots, the same character or subject holds up visually from one cut to the next, without extra prompting to keep it consistent.
Multi-shot sequences, clean on-screen text
Generate a full sequence with hard cuts inside one call instead of stitching shots together, and render titles or animated typography cleanly right in the frame.
Built for length and resolution
Generate clips from 5 to 20 seconds at 480p or 1080p, across aspect ratios spanning cinematic 21:9 to vertical 9:21.
Built for real workflows
FLUX 3 Video is Black Forest Labs' first step into video, built from a unified multimodal architecture rather than an image model stretched to fit. It covers the full stylistic range in a single model, from candid camcorder footage to cinematic photoreal, so switching styles doesn't mean switching tools.
Product and brand footage
Turn a single product photo into polished commercial footage, with motion and sound generated to match, and no live shoot required.
Multi-shot cinematic sequences
Generate a full sequence with hard cuts inside one call, so establishing shots and follow-ups come back together with audio that matches the action.
Motion design and on-screen text
Produce stylized motion graphics with clean in-video titles and animated typography, from the same model used for photoreal footage.
Storyboarding with keyframes
Pin the key moments of a shot and let the model generate everything in between, closer to directing than prompting.
Localized dialogue and voiceover
Generate native multilingual dialogue synced to the action, so the same clip can carry a script in more than one language without a separate dubbing pass.
Stylized and handmade looks
Move between candid camcorder footage, animation, and motion design from the same model, so a campaign can shift styles without switching tools.
Frequently asked questions
What is FLUX 3 Video?
Black Forest Labs' multimodal foundation model for video generation with synchronized audio, trained jointly across image, video, audio, and action. It covers text to video and image to video on one architecture, with keyframe control layered on top to pin what happens on screen.
What generation modes does it support?
Text to video and image to video, plus keyframe control to pin specific frames in a sequence. All of it runs through the same Runware integration.
Does it generate audio automatically?
Yes, by default across every mode, and it can be switched off per request. Audio comes from the same joint training that grounds the model's sense of motion, rather than a separate pass added afterward.
What resolutions and durations are supported?
Clips from 5 to 20 seconds at 480p or 1080p, across seven aspect ratios: 21:9, 16:9, 4:3, 1:1, 3:4, 9:16, and 9:21.
Do I need separate integrations for each mode?
No. Text to video and image to video, plus keyframe control, all run through the same Runware API as every other model on the platform, under the same key and billing.
How much does FLUX 3 Video cost?
Pricing is per second of output rather than a flat per-clip rate. Text to video and image to video cost $0.17 per second at 720p and $0.29 per second at 1080p. Keyframe control runs on the same per-second, usage-based pricing.
Is FLUX 3 Video available now?
Yes. FLUX 3 Video is live on Runware, available through the API and Playground under the same key and billing as every other model on the platform.
Talk to us about production volume
Have questions about FLUX 3 Video? Chat to our team about enterprise usage, including volume discounts and dedicated RPM, and we'll follow up shortly.