Gemini Omni Flash 1.1

Gemini Omni Flash 1.1 is Google's updated multimodal video generation and editing model in the Gemini Omni family. It generates native synchronized audio from the prompt, and expands the original Omni Flash workflow with additional 1080p and 4K output modes, a 360p draft mode, scene extension in 3 to 10 second increments up to 30 seconds total, start-to-end frame interpolation for fluid transitions, and reference-to-video generation guided by both images and short video clips. It is built for teams that need stronger continuity, higher-resolution delivery, and more controllable multi-input video creation than the first Omni Flash release.

Complete technical specification for integration
Step-by-step tutorials for advanced use cases
← All GuidesResolution and duration in Gemini Omni Flash 1.1
How to size and time a Gemini Omni Flash 1.1 video: the four resolution presets, the supported width and height pairs, and how tier and duration drive what a clip costs.
Introduction
Gemini Omni Flash 1.1 delivers the same shot across four tiers, from a 360p draft to a 4K master, and the tier you pick is the single biggest lever on what a clip costs. A 4K second is roughly nine times a 360p second, so the tier decision belongs at the start of a workflow rather than at the end of one.
Duration is the other half of the same equation. The model produces 3 to 10 seconds per call and bills by the second, which makes tier and runtime the two numbers worth deciding before you write the prompt.
A slow macro orbit around a single premium leather sneaker on a brushed concrete plinth, in a single unbroken scene. Raking side light picks out the perforation pattern on the toe box, the waxed contrast stitching along the sole, and the fine grain of the white leather upper. E-commerce packshot cinematography, shallow depth of field, seamless pale grey backdrop. The audio is a quiet studio room tone, no music.
This guide covers the two ways to size a clip, when each one applies, what the tiers actually buy you, and how to pick a duration.
Two ways to size a clip
Sizing goes through either resolution or a width/height pair, and the request is rejected if you send both. They cover the same output sizes from different directions.
import { createClient } from '@runware/sdk'
const client = await createClient({ apiKey: process.env.RUNWARE_API_KEY })
await client.connect()
const [result] = await client.run({
model: 'google:gemini@omni-flash-1.1',
positivePrompt: 'A slow push in on a hiker resting against a boulder on an exposed alpine ridge in the late afternoon, in a single unbroken scene. Outdoor apparel brand cinematography. Low warm sun raking across lichen-covered rock. The audio is steady wind across the ridge, no music.',
resolution: '1080p',
duration: 6
})import asyncio
import os
from runware import Runware
async def main():
async with Runware(api_key=os.environ["RUNWARE_API_KEY"]) as client:
results = await client.run({
"model": "google:gemini@omni-flash-1.1",
"positivePrompt": "A slow push in on a hiker resting against a boulder on an exposed alpine ridge in the late afternoon, in a single unbroken scene. Outdoor apparel brand cinematography. Low warm sun raking across lichen-covered rock. The audio is steady wind across the ridge, no music.",
"resolution": "1080p",
"duration": 6
})
asyncio.run(main())curl https://api.runware.ai/v1 \
-H "Authorization: Bearer $RUNWARE_API_KEY" \
-H "Content-Type: application/json" \
-d '[
{
"taskType": "videoInference",
"taskUUID": "c8d3e2f5-9a0b-4134-c2d3-e4f506172839",
"model": "google:gemini@omni-flash-1.1",
"positivePrompt": "A slow push in on a hiker resting against a boulder on an exposed alpine ridge in the late afternoon, in a single unbroken scene. Outdoor apparel brand cinematography. Low warm sun raking across lichen-covered rock. The audio is steady wind across the ridge, no music.",
"resolution": "1080p",
"duration": 6
}
]'runware run google:gemini@omni-flash-1.1 \
positivePrompt="A slow push in on a hiker resting against a boulder on an exposed alpine ridge in the late afternoon, in a single unbroken scene. Outdoor apparel brand cinematography. Low warm sun raking across lichen-covered rock. The audio is steady wind across the ridge, no music." \
resolution=1080p \
duration=6{
"taskType": "videoInference",
"taskUUID": "c8d3e2f5-9a0b-4134-c2d3-e4f506172839",
"model": "google:gemini@omni-flash-1.1",
"positivePrompt": "A slow push in on a hiker resting against a boulder on an exposed alpine ridge in the late afternoon, in a single unbroken scene. Outdoor apparel brand cinematography. Low warm sun raking across lichen-covered rock. The audio is steady wind across the ridge, no music.",
"resolution": "1080p",
"duration": 6
}[
{
"taskType": "videoInference",
"taskUUID": "c8d3e2f5-9a0b-4134-c2d3-e4f506172839",
"videoUUID": "5b6c7d8e-9f0a-4123-b4c5-d6e7f8091234",
"videoURL": "https://vm.runware.ai/video/os/a14d18/ws/2/vi/5b6c7d8e-9f0a-4123-b4c5-d6e7f8091234.mp4"
}
]resolution takes 360p, 720p, 1080p, or 4K and names a tier without committing to an orientation. When the call also carries input media, the tier follows the aspect ratio of that media, which is what makes it the right choice for editing and extension work.
width and height travel together, both or neither, and must match one of eight supported pairs:
| Tier | 16:9 | 9:16 |
|---|---|---|
| 360p | 640 × 360 | 360 × 640 |
| 720p | 1280 × 720 | 720 × 1280 |
| 1080p | 1920 × 1080 | 1080 × 1920 |
| 4K | 3840 × 2160 | 2160 × 3840 |
Landscape and portrait are the only orientations. There is no square output and no arbitrary pair, so a value off this table fails validation rather than snapping to the nearest legal size.
Portrait is a first-class output rather than a crop of a landscape render, which matters for social work where the subject has to sit in a tall frame from the start.
A vertical shot of a hand lifting a chilled bottle of iced tea from a lit retail cooler shelf, in a single unbroken scene. Framed tall for a social story with the bottle centred and room above and below. Condensation on the glass, neat rows of identical bottles behind it. Bright clean retail lighting. Beverage brand social cinematography. The audio is the soft suction of the cooler door and a quiet store ambience, no music.
What each tier buys
The four tiers below ran the same prompt, one call each. Because the model exposes no seed, these are four separate takes rather than one take rendered four ways, so read them for fidelity rather than for a frame-by-frame match.
A slow push in on a hiker resting against a boulder on an exposed alpine ridge in the late afternoon, in a single unbroken scene. She wears a technical shell jacket in burnt orange with a visible woven texture and taped seams, a coiled rope and a worn leather boot in the foreground. Outdoor apparel brand cinematography, crisp and natural. Low warm sun raking across lichen-covered rock and dry grass. The audio is steady wind across the ridge and the faint creak of the jacket fabric, no music.
A slow push in on a hiker resting against a boulder on an exposed alpine ridge in the late afternoon, in a single unbroken scene. She wears a technical shell jacket in burnt orange with a visible woven texture and taped seams, a coiled rope and a worn leather boot in the foreground. Outdoor apparel brand cinematography, crisp and natural. Low warm sun raking across lichen-covered rock and dry grass. The audio is steady wind across the ridge and the faint creak of the jacket fabric, no music.
A slow push in on a hiker resting against a boulder on an exposed alpine ridge in the late afternoon, in a single unbroken scene. She wears a technical shell jacket in burnt orange with a visible woven texture and taped seams, a coiled rope and a worn leather boot in the foreground. Outdoor apparel brand cinematography, crisp and natural. Low warm sun raking across lichen-covered rock and dry grass. The audio is steady wind across the ridge and the faint creak of the jacket fabric, no music.
A slow push in on a hiker resting against a boulder on an exposed alpine ridge in the late afternoon, in a single unbroken scene. She wears a technical shell jacket in burnt orange with a visible woven texture and taped seams, a coiled rope and a worn leather boot in the foreground. Outdoor apparel brand cinematography, crisp and natural. Low warm sun raking across lichen-covered rock and dry grass. The audio is steady wind across the ridge and the faint creak of the jacket fabric, no music.
Web delivery flattens most of that difference, because every clip on this page is re-encoded for the browser. The pair below is the honest version: one frame pulled from the 360p take and one from the 4K take, cropped to the same region and shown at their own pixel density.


The fabric weave and the detail in her face survive the crop at 4K and dissolve into flat colour at 360p. That gap is the whole argument for the tier ladder: 360p is for judging a take, not for judging a texture.
Drafting cheap, finishing expensive
Per second of output, the tiers price at $0.036 for 360p, $0.10 for 720p, $0.16 for 1080p, and $0.32 for 4K. A six-second clip therefore runs about $0.22 at the draft tier against $1.92 at the master tier.
That spread is what makes an iterate-then-commit loop worth building. Run prompt variations at 360p until the framing, the action, and the audio are right, then re-run the winning prompt once at the delivery tier. Ten drafts plus one 4K master costs less than three 4K attempts and gets you more looks at the idea.
A 360p draft is a genuinely different take from the 4K run of the same prompt, so treat drafting as a way to settle prompt wording rather than to lock a specific take. What survives the tier change is everything the prompt says. What does not survive is the incidental staging the model chose on that particular run. Drop settings.temperature if you want the two runs to land closer together, as covered in the prompting guide.
Choosing a duration
duration accepts a whole number of seconds from 3 to 10 and defaults to 6. It is not a quality control, it is a container: the model paces whatever you asked for to fill the time you gave it.
A single quick beat: a cordless vacuum clicks into its wall dock and the charge light turns green, in a single unbroken scene. Close medium shot on the dock against a pale hallway wall, bright even daylight. Home appliance commercial cinematography. The audio is the firm plastic click of the dock and a short soft chime, no music.
A continuous shot of a florist assembling a bouquet at a workbench from first stem to finished wrap, in a single unbroken scene with no cuts. She lays in eucalyptus, adds three cream roses, turns the bunch as she binds it with twine, then folds brown paper around it and sets it upright. Steady medium shot from across the bench, soft daylight from a shopfront window. Small-business retail cinematography. The audio is rustling paper, snipping shears, and a quiet shop ambience, no music.
The three-second clip carries one action and ends on it, which is the right shape for a product-detail loop or a UI transition. The ten-second clip carries a sequence with a beginning and an end, and the florist prompt names four steps that the model has room to pace out.
Ask for four actions in three seconds and each one gets under a second, which reads as rushed. Ask for one action in ten seconds and the model pads, usually with camera drift. Match the number of beats in the prompt to the seconds you bought.
To go past ten seconds, generate a clip and extend it. See the extension guide.
Tips
-
Pick
resolutionorwidth/height, never both. Sending both fails validation. The preset is simpler when orientation is already implied, the pair is explicit when it is not. -
Use
resolutionfor anything with input media. It takes its aspect ratio from the source, andwidth/heightare rejected alongsideinputs.videoanyway. -
Draft at 360p. At roughly a ninth the per-second cost of 4K, the draft tier is where prompt wording gets settled. Commit to a delivery tier once.
-
Don't judge texture from a draft. Fabric, skin, foliage, and fine type all collapse at 360p. Judge framing, action, and audio there, and judge detail at 1080p or above.
-
Generate portrait natively. The 9:16 pairs are real output sizes, and framing a subject for a tall frame beats cropping a landscape render later.
-
Match beats to seconds. One action fits 3 seconds, a short sequence fits 10. Overstuffing a short clip rushes it and understuffing a long one invites drift.
-
Remember the 10-second ceiling. It is the maximum for a single call, and longer pieces come from extending a clip rather than from a bigger
duration.