Best Image-to-Video
Models selected for turning still images into short video clips with coherent motion and stable subjects. Useful for simple animation, camera movement, and bringing static visuals to life.
The picks.Animate static images
MiniMax H3
Multimodal video generation with native synced audio, multi-reference consistency, and continuation workflows
Gemini Omni Flash 1.1
4K multimodal video generation and editing with native audio, clip extension, start-to-end interpolation, and reference video control
Wan3.0
All-in-one multimodal video generation with native 30-second clips, large reference capacity, and precise video editing
PixVerse V6
Multi-shot cinematic video generation with native audio, 20+ camera controls, and character consistency
Seedance 2.0
Unified multimodal audio-video generation with multi-reference input and physics-aware motion
HappyHorse 1.1
Multimodal video generation with stronger motion expressiveness, multi-image reference consistency, and improved audio-visual sync
Grok Imagine Video 1.5
Higher-tier Grok image-to-video generation from a single starting frame with longer durations and stronger output quality
SkyReels V4
Unified multimodal video model for generation, inpainting, and editing with synchronized audio
Grok Imagine Video
AI video generation with synchronized audio from text and images
Kling VIDEO 3.0 Omni Pro
Unified multimodal video generation with native audio and higher-fidelity renders
HappyHorse-1.0
Text-to-video and image-to-video model with 1080p output and short-form clip control
Kling VIDEO 3.0 Standard
Multimodal video generation with native audio and efficient performance
Wan2.7
Multimodal video generation with reference consistency, video editing, and native audio
Kling VIDEO 3.0 Omni Standard
Cost-efficient multimodal video generation with native audio and editing
Kling VIDEO 2.6 Standard
High-quality AI video generation with strong motion and camera control
MiniMax H3 Max Turbo
Distilled H3 Max video generation for faster iteration with native audio and first-to-last frame control
MiniMax H3 Fast
Fast multimodal video generation with native audio and image, video, and audio reference control
Common questions
MiniMax H3 Fast, released September 2026 per the live catalog. Membership updates automatically as the catalog publishes new models to this collection.
Production notes
Every price on this page is the model's published rate from the live Runware catalog, using the cheapest listed configuration unless stated otherwise. Prices vary with resolution, duration, quality tier, or token volume, so check the pricing table on each model page before estimating unit economics.
The Runware catalog does not publish per-model latency figures, so this page does not quote end-to-end timings. Where a model's own description commits to speed (for example sub-second generation or realtime streaming), that claim is repeated here. For anything else, benchmark the exact models in the Playground with your own payload sizes before committing to an SLA.
Models are addressed by versioned AIR identifiers, so a workflow pinned to specific model versions keeps producing the same behaviour as new versions ship. Adopt upgrades deliberately by re-running your evaluation set against the new version before switching production traffic.
About this collection
Models selected for turning still images into short video clips with coherent motion and stable subjects. Useful for simple animation, camera movement, and bringing static visuals to life.
Membership comes directly from the Runware catalog: the 20 models on this page are the live catalog's own membership for the "Best Image-to-Video" collection. Names, descriptions, pricing, capability chips, samples, and guides are all read live from the catalog — nothing here is hand-curated.
Model guides.Learn how to use the stack
Editing video
How to edit a finished clip with a prompt in MiniMax H3: replace a subject, relight a scene, add or remove elements, and combine edits, keeping the rest untouched.
Read the guide →First and last frame
How to animate a still image with MiniMax H3 and bridge a first and last frame into one continuous shot, with the output following the image's own aspect ratio.
Read the guide →Generating video
How to generate video from text with MiniMax H3: the six 2K aspect ratios, 5 to 15 second durations, native synced audio, and prompting for cinematic shots.
Read the guide →Motion, camera, and performance
How to transfer motion, a camera move, or an acting performance from a reference video onto a new subject with MiniMax H3 Omni Reference.
Read the guide →Reference-driven consistency
How to lock a character, product, or style across a new MiniMax H3 shot with Omni Reference images, and address each reference by index in the prompt.
Read the guide →Sound and voice
How to direct MiniMax H3's native audio from the prompt: ambience, synced sound effects, music, and spoken dialogue with lip-sync.
Read the guide →Editing video
How to edit an existing clip with Gemini Omni Flash 1.1: relighting, weather, and restyling through inputs.video, why short prompts win, and how to chain edits.
Read the guide →Extending a video
How to extend a clip past the 10-second ceiling with Gemini Omni Flash 1.1: sending duration with inputs.video, prompting continuation, and chaining extensions.
Read the guide →First and last frame control
How to pin the opening and closing frames of a Gemini Omni Flash 1.1 clip with inputs.frameImages, and prompt the transition the model builds between them.
Read the guide →Prompting
How to prompt Gemini Omni Flash 1.1 for video: the five-element structure, camera vocabulary, holding a single shot, directing audio, and the sampling controls.
Read the guide →Reference images and videos
How to carry a character, product, or style into a Gemini Omni Flash 1.1 shot with inputs.referenceImages and inputs.referenceVideos, and when to combine the two.
Read the guide →Resolution and duration
How to size and time a Gemini Omni Flash 1.1 video: the four resolution presets, the supported width and height pairs, and how tier and duration drive what a clip costs.
Read the guide →Document and web inputs with Wan 3.0
How to turn structured content into video with Wan 3.0: recipes as cooking videos and public webpages as short landmark or historical montages.
Read the guide →Image inputs with Wan 3.0
How to use images with Wan 3.0's video generation: pin an image as the first frame, morph between two frames, and carry subject or product identity across new scenes with reference images.
Read the guide →Prompting Wan 3.0
How to write prompts for Wan 3.0 that get you the shot: how the model reads layered directives, camera language, style and register range, directing native audio, and the dimensions and duration rules.
Read the guide →Video and audio references with Wan 3.0
How to use video and audio references with Wan 3.0: carrying a subject across new scenes with referenceVideos, driving video from an audio source with referenceAudios, and combining multiple reference types in one call.
Read the guide →Multi-shot storytelling with anchor frames
How to use PixVerse V6 to generate multi-shot video reels from one or two anchor frames in a single call, using the beat-structured prompt pattern.
Read the guide →Cinematic shot direction with storyboard prompts
How to write multi-shot storyboard prompts for HappyHorse 1.1 to direct cinematic sequences with shot sizes, camera movement, and subject continuity in a single call.
Read the guide →Casting multiple characters with reference images
How to use HappyHorse 1.1's reference workflow to cast one or more characters into a generated video and preserve their identity through every cut.
Read the guide →