Best Image-to-Video
Models selected for turning still images into short video clips with coherent motion and stable subjects. Useful for simple animation, camera movement, and bringing static visuals to life.
The picks.Animate static images
MiniMax H3
Multimodal video generation with native synced audio, multi-reference consistency, and continuation workflows
Gemini Omni Flash
Multimodal video generation and editing with native audio, multi-turn control, and photo-to-video references
Seedance 2.0
Unified multimodal audio-video generation with multi-reference input and physics-aware motion
Seedance 2.5
Professional multimodal video generation with native 30-second clips, large-scale reference control, and precise localized editing
Grok Imagine Video 1.5
Higher-tier Grok image-to-video generation from a single starting frame with longer durations and stronger output quality
PixVerse V6
Multi-shot cinematic video generation with native audio, 20+ camera controls, and character consistency
HappyHorse 1.1
Multimodal video generation with stronger motion expressiveness, multi-image reference consistency, and improved audio-visual sync
SkyReels V4
Unified multimodal video model for generation, inpainting, and editing with synchronized audio
Kling VIDEO 3.0 Omni Pro
Unified multimodal video generation with native audio and higher-fidelity renders
Kling VIDEO 3.0 Pro
High-fidelity multimodal video generation with native audio and advanced editing
Kling VIDEO 2.6 Standard
High-quality AI video generation with strong motion and camera control
Wan2.7
Multimodal video generation with reference consistency, video editing, and native audio
Grok Imagine Video
AI video generation with synchronized audio from text and images
Kling VIDEO 3.0 Omni Standard
Cost-efficient multimodal video generation with native audio and editing
Kling VIDEO 3.0 Standard
Multimodal video generation with native audio and efficient performance
PixVerse V5.6
Enhanced cinematic video generation with improved lip-sync and audio realism
HappyHorse-1.0
Text-to-video and image-to-video model with 1080p output and short-form clip control
Kling VIDEO 2.6 Pro
Kling VIDEO 2.6 Pro is a full audio-visual AI video model that combines cinematic-quality video generation with native audio (dialogue, sound effects, ambience), with optional Motion Control for precise character movement via the API.
Common questions
Seedance 2.5, released August 2026 per the live catalog. Membership updates automatically as the catalog publishes new models to this collection.
Production notes
Every price on this page is the model's published rate from the live Runware catalog, using the cheapest listed configuration unless stated otherwise. Prices vary with resolution, duration, quality tier, or token volume, so check the pricing table on each model page before estimating unit economics.
The Runware catalog does not publish per-model latency figures, so this page does not quote end-to-end timings. Where a model's own description commits to speed (for example sub-second generation or realtime streaming), that claim is repeated here. For anything else, benchmark the exact models in the Playground with your own payload sizes before committing to an SLA.
Models are addressed by versioned AIR identifiers, so a workflow pinned to specific model versions keeps producing the same behaviour as new versions ship. Adopt upgrades deliberately by re-running your evaluation set against the new version before switching production traffic.
About this collection
Models selected for turning still images into short video clips with coherent motion and stable subjects. Useful for simple animation, camera movement, and bringing static visuals to life.
Membership comes directly from the Runware catalog: the 20 models on this page are the live catalog's own membership for the "Best Image-to-Video" collection. Names, descriptions, pricing, capability chips, samples, and guides are all read live from the catalog — nothing here is hand-curated.
Model guides.Learn how to use the stack
Editing video
How to edit a finished clip with a prompt in MiniMax H3: replace a subject, relight a scene, add or remove elements, and combine edits, keeping the rest untouched.
Read the guide →First and last frame
How to animate a still image with MiniMax H3 and bridge a first and last frame into one continuous shot, with the output following the image's own aspect ratio.
Read the guide →Generating video
How to generate video from text with MiniMax H3: the six 2K aspect ratios, 5 to 15 second durations, native synced audio, and prompting for cinematic shots.
Read the guide →Motion, camera, and performance
How to transfer motion, a camera move, or an acting performance from a reference video onto a new subject with MiniMax H3 Omni Reference.
Read the guide →Reference-driven consistency
How to lock a character, product, or style across a new MiniMax H3 shot with Omni Reference images, and address each reference by index in the prompt.
Read the guide →Sound and voice
How to direct MiniMax H3's native audio from the prompt: ambience, synced sound effects, music, and spoken dialogue with lip-sync.
Read the guide →Cinematic prompting
How to prompt Gemini Omni Flash for cinematic video using Google's five-element structure, camera language, and the less-prescriptive sweet spot.
Read the guide →Editing video
How to edit existing footage with Gemini Omni Flash's inputs.video parameter to relight, restyle, swap weather, or add characters while preserving the source's composition and motion.
Read the guide →Reference-driven video
How to use Gemini Omni Flash's reference image workflow to lock a visual style, hold a character across scenes, or guide a video through storyboard key beats.
Read the guide →Editing video
How to edit an existing clip in place with Seedance 2.5 using inputs.video: remove and replace elements, restyle and relight the frame, swap backgrounds, and edit by timestamp.
Read the guide →Image to video and keyframes
How to turn a still into video with Seedance 2.5: animate a first frame, interpolate a first and last frame, loop a still, and sequence images as keyframes.
Read the guide →Long-form and extension
How to structure a 30-second Seedance 2.5 brief as one continuous take or a multi-shot passage, hold continuity across the runtime, and extend past 30 seconds with inputs.video.
Read the guide →Motion and performance transfer
How to drive Seedance 2.5 with a reference video for motion, camera, and lip-sync while a reference image supplies identity: motion transfer, character swaps, and clay to finished shot.
Read the guide →Multimodal reference
How to compose up to 30 images, 10 videos, and 10 audio clips into one Seedance 2.5 call with typed @-tag addressing, from cast ensembles and product families to audio-driven scenes.
Read the guide →Multilingual video
How to generate video in 10+ languages with Seedance 2.5: native lip-sync from a text prompt, switching languages within one clip, and controlling accent and delivery.
Read the guide →Prompting
How to prompt Seedance 2.5 for directed text-to-video: the five-layer shot scaffold, camera vocabulary the model reads directly, second-level timing, and native audio.
Read the guide →Multi-shot storytelling with anchor frames
How to use PixVerse V6 to generate multi-shot video reels from one or two anchor frames in a single call, using the beat-structured prompt pattern.
Read the guide →Cinematic shot direction with storyboard prompts
How to write multi-shot storyboard prompts for HappyHorse 1.1 to direct cinematic sequences with shot sizes, camera movement, and subject continuity in a single call.
Read the guide →Casting multiple characters with reference images
How to use HappyHorse 1.1's reference workflow to cast one or more characters into a generated video and preserve their identity through every cut.
Read the guide →Directing motion in image-to-video prompts
How to write Runway Gen-4.5 image-to-video prompts that direct motion instead of redescribing the scene, using the camera and subject channels.
Read the guide →