MiniMax H3

Live now on Runware

MiniMax's omni-modal video model with native synced audio and reference-guided editing

Access MiniMax H3, MiniMax's general-purpose video model that generates and edits clips from text, image, video, and audio inputs, with native stereo audio in every output.

0:00

All workflows supported natively.

Text-to-videoKeyframe-guided videoReference-to-videoInstruction-based editing

One integration for generation and editing, with no separate endpoints.

Four ways to build a video

MiniMax H3 can generate an entirely new clip from a written prompt, follow frames pinned to specific points in a sequence, carry a subject in from reference inputs, or edit a clip you already have. All of it runs through the same Runware integration.

Text to video

A prompt alone is enough. MiniMax H3 generates an entirely new clip from scratch, with synced audio timed to the action as it happens.

Keyframe-guided video

Pin a first frame, a last frame, or both, and H3 fills in the motion that connects them. It reads closer to storyboarding than prompting.

Multi-reference conditioning

Combine reference images, video, and audio in a single call to carry a subject's look, motion, or voice into a brand new scene.

Instruction-based editing

Point at an existing clip, describe the change as an instruction, and get back the same shot with just that change applied.

Why MiniMax H3 stands out

Native audio, generated by default

Every clip returns with synchronized stereo audio grounded in the action on screen, timed to specific moments like a hammer strike or a footstep rather than layered on as generic ambience.

Keyframe-guided generation

Pin a first frame, a last frame, or both, and H3 generates the motion that connects them, closer to storyboarding than prompting.

Multi-reference conditioning

Combine reference images, video, and audio in one call to keep a subject's look, motion, or voice consistent as it moves into a new scene.

Instruction-based editing

Point at an existing clip, describe the change as an instruction, and get back the same shot with just that change applied, while the subject and camera path hold.

2K across six aspect ratios

Renders at a fixed set of 2K size pairs spanning cinematic 21:9 to vertical 9:16, so the frame is chosen before generation instead of cropped after.

Directed with cinematic language

Responds to camera, lens, and lighting terms like low-angle dolly-in or chiaroscuro lighting, so a shot can be directed rather than just described.

See it in action

Real generations and edits from MiniMax H3, spanning a fresh clip built from an image and three instruction-based edits made to existing footage.

Seed patent documentary title teaser

Image to video

Maritime museum scene replacement

Edit

Robotics team awards portrait

Edit

Restored deaf pottery workshop film

Edit

Built for real workflows

MiniMax H3 is a general-purpose, omni-modal video model: unified understanding across text, images, video, and audio, with generation and editing running through the same architecture. It's built for production work that starts from a prompt, a reference, or a clip you already have.

Product and brand footage

Swap an object, relight a scene, or add a prop to existing footage without a reshoot.

Reference-driven character performance

Carry a subject's look, motion, or voice into a brand new scene using a mix of image, video, and audio references.

Storyboarding with keyframes

Pin the first and last frame of a shot and let H3 generate the motion in between.

Clip continuation

Extend an existing clip or audio segment into a seamless new video, rather than starting the shot over.

Scene relighting and background swaps

Change the time of day or the backdrop while the subject, geometry, and camera path stay put.

Social-ready vertical video

Generate directly at 2K in six formats, from cinematic 21:9 to vertical 9:16, without cropping after the fact.

Frequently asked questions

What is MiniMax H3?

MiniMax's general-purpose, omni-modal video model. It generates and edits video from text, image, video, and audio inputs, with native synchronized stereo audio in every output rather than a separate dubbing pass.

What generation modes does it support?

Text to video, first-frame and keyframe-guided generation, multi-reference conditioning that combines reference images, video, and audio in one call, and instruction-based editing of an existing clip. All of it runs through the same Runware integration.

Does it generate audio automatically?

Yes, by default. Every clip comes back with synced stereo sound generated alongside the picture, and it responds to a written Sound: clause the same way it responds to a description of the image.

What resolutions and durations are supported?

Clips from 5 to 15 seconds at 2K, across six fixed aspect ratios: 16:9 at 2560x1440, 21:9 at 2944x1248, 4:3 at 1920x1440, 1:1 at 1440x1440, 3:4 at 1440x1920, and 9:16 at 1440x2560.

How does editing work?

Pass an existing clip in referenceVideos, then write the prompt as an instruction rather than a scene description, naming the change and what should stay untouched. MiniMax H3 returns the same shot with just that change applied, holding the subject, motion, and camera path.

Do I need separate integrations for each mode?

No. Generation, multi-reference composition, and editing all run through the same Runware API endpoint, under the same key and billing as every other model on the platform.

How much does MiniMax H3 cost?

Pricing is metered per second of output rather than a flat per-clip rate, at $0.1391 per second of 2K video. A 5-second clip, the shortest supported length, costs $0.6955.

Is MiniMax H3 available now?

Yes. MiniMax H3 is live on Runware, available through the API and Playground under the same key and billing as every other model on the platform.

Talk to us about volume discounts

Have questions about MiniMax H3? Chat to our team about enterprise usage, including volume discounts and dedicated RPM, and we'll follow up shortly.