HeyGen Video 1.0

HeyGen Video 1.0 is HeyGen's video generation model for short clips with synchronized audio. It generates the full frame from a text prompt, a first-frame image, or up to 12 image, video, and audio references, and produces dialogue, ambience, and sound effects in the same pass, with no separate lip-sync step. It outputs 5 to 15 second clips at 480p or 768p, suited to presenter shots, product and training content, and social video.

Complete technical specification for integration
Ready-to-use code snippets for common workflows
Step-by-step tutorials for advanced use cases
← All GuidesDialogue, ambience and sound effects
How to direct HeyGen Video 1.0's sound: dialogue sized to the clip length, two speakers, other languages, ambience by distance, effects on cue and music beds.
Introduction
HeyGen Video 1.0 generates the soundtrack in the same pass as the picture. Dialogue, room tone and sound effects all come out of the one prompt, with the lips animated to the speech in that same pass. Because the sound and the image are built together, the sound matches the space the prompt describes: a glass greenhouse sounds different from an empty bedroom.
Every clip comes back with an AAC stereo track, and what goes on it is up to the prompt. The garden center owner below speaks a line written word for word, over the greenhouse sounds that the prompt also named.
Shot on a large-format cinema camera with a 50mm lens at T2.8, static at chest height, a waist-up medium shot that shows her whole head with a little space above her hair. The owner of a small garden center, a woman in her fifties with short curly gray hair and a plain olive linen shirt, stands between long tables of potted herbs in a bright greenhouse, facing the camera and holding a small potted rosemary in both hands. She lifts the pot slightly toward the camera and says, in a warm, practical voice: "Water deeply once a week instead of a little every day. The roots follow the water down, and the plant grows tougher." Her lips stay in sync with every word. Soft daylight through the greenhouse glass, rows of plants softly out of focus behind her. Real skin texture, no retouching, no beauty filter. No logos, brand names or printed words anywhere in frame, including on the pots. Audio: her voice close and clear, the soft hiss of a sprinkler somewhere behind her, and birdsong outside the glass. No music.
This guide covers how the model treats sound you do not mention, how to write dialogue and size it to the clip, two speakers and other languages, and how to direct ambience, effects and music.
The request
There is no audio parameter. Sound direction lives in positivePrompt, and duration decides how much dialogue fits.
import { createClient } from '@runware/sdk'
const client = await createClient({ apiKey: process.env.RUNWARE_API_KEY })
await client.connect()
const [result] = await client.run({
model: 'heygen:video@1.0',
positivePrompt: 'Shot on a large-format cinema camera with a 50mm lens at T2.8, static at chest height, a waist-up medium shot that shows her whole head with a little space above her hair. The owner of a small garden center, a woman in her fifties with short curly gray hair and a plain olive linen shirt, stands between long tables of potted herbs in a bright greenhouse, facing the camera and holding a small potted rosemary in both hands. She lifts the pot slightly toward the camera and says, in a warm, practical voice: "Water deeply once a week instead of a little every day. The roots follow the water down, and the plant grows tougher." Her lips stay in sync with every word. Soft daylight through the greenhouse glass, rows of plants softly out of focus behind her. Real skin texture, no retouching, no beauty filter. No logos, brand names or printed words anywhere in frame, including on the pots. Audio: her voice close and clear, the soft hiss of a sprinkler somewhere behind her, and birdsong outside the glass. No music.',
width: 1344,
height: 768,
duration: 11
})import asyncio
import os
from runware import Runware
async def main():
async with Runware(api_key=os.environ["RUNWARE_API_KEY"]) as client:
results = await client.run({
"model": "heygen:video@1.0",
"positivePrompt": "Shot on a large-format cinema camera with a 50mm lens at T2.8, static at chest height, a waist-up medium shot that shows her whole head with a little space above her hair. The owner of a small garden center, a woman in her fifties with short curly gray hair and a plain olive linen shirt, stands between long tables of potted herbs in a bright greenhouse, facing the camera and holding a small potted rosemary in both hands. She lifts the pot slightly toward the camera and says, in a warm, practical voice: \"Water deeply once a week instead of a little every day. The roots follow the water down, and the plant grows tougher.\" Her lips stay in sync with every word. Soft daylight through the greenhouse glass, rows of plants softly out of focus behind her. Real skin texture, no retouching, no beauty filter. No logos, brand names or printed words anywhere in frame, including on the pots. Audio: her voice close and clear, the soft hiss of a sprinkler somewhere behind her, and birdsong outside the glass. No music.",
"width": 1344,
"height": 768,
"duration": 11
})
asyncio.run(main())curl https://api.runware.ai/v1 \
-H "Authorization: Bearer $RUNWARE_API_KEY" \
-H "Content-Type: application/json" \
-d '[
{
"taskType": "videoInference",
"taskUUID": "9b2e5f71-c8a4-4d13-b6f0-2a7e1d94c358",
"model": "heygen:video@1.0",
"positivePrompt": "Shot on a large-format cinema camera with a 50mm lens at T2.8, static at chest height, a waist-up medium shot that shows her whole head with a little space above her hair. The owner of a small garden center, a woman in her fifties with short curly gray hair and a plain olive linen shirt, stands between long tables of potted herbs in a bright greenhouse, facing the camera and holding a small potted rosemary in both hands. She lifts the pot slightly toward the camera and says, in a warm, practical voice: \"Water deeply once a week instead of a little every day. The roots follow the water down, and the plant grows tougher.\" Her lips stay in sync with every word. Soft daylight through the greenhouse glass, rows of plants softly out of focus behind her. Real skin texture, no retouching, no beauty filter. No logos, brand names or printed words anywhere in frame, including on the pots. Audio: her voice close and clear, the soft hiss of a sprinkler somewhere behind her, and birdsong outside the glass. No music.",
"width": 1344,
"height": 768,
"duration": 11
}
]'runware run heygen:video@1.0 \
positivePrompt="Shot on a large-format cinema camera with a 50mm lens at T2.8, static at chest height, a waist-up medium shot that shows her whole head with a little space above her hair. The owner of a small garden center, a woman in her fifties with short curly gray hair and a plain olive linen shirt, stands between long tables of potted herbs in a bright greenhouse, facing the camera and holding a small potted rosemary in both hands. She lifts the pot slightly toward the camera and says, in a warm, practical voice: \"Water deeply once a week instead of a little every day. The roots follow the water down, and the plant grows tougher.\" Her lips stay in sync with every word. Soft daylight through the greenhouse glass, rows of plants softly out of focus behind her. Real skin texture, no retouching, no beauty filter. No logos, brand names or printed words anywhere in frame, including on the pots. Audio: her voice close and clear, the soft hiss of a sprinkler somewhere behind her, and birdsong outside the glass. No music." \
width=1344 \
height=768 \
duration=11{
"taskType": "videoInference",
"taskUUID": "9b2e5f71-c8a4-4d13-b6f0-2a7e1d94c358",
"model": "heygen:video@1.0",
"positivePrompt": "Shot on a large-format cinema camera with a 50mm lens at T2.8, static at chest height, a waist-up medium shot that shows her whole head with a little space above her hair. The owner of a small garden center, a woman in her fifties with short curly gray hair and a plain olive linen shirt, stands between long tables of potted herbs in a bright greenhouse, facing the camera and holding a small potted rosemary in both hands. She lifts the pot slightly toward the camera and says, in a warm, practical voice: \"Water deeply once a week instead of a little every day. The roots follow the water down, and the plant grows tougher.\" Her lips stay in sync with every word. Soft daylight through the greenhouse glass, rows of plants softly out of focus behind her. Real skin texture, no retouching, no beauty filter. No logos, brand names or printed words anywhere in frame, including on the pots. Audio: her voice close and clear, the soft hiss of a sprinkler somewhere behind her, and birdsong outside the glass. No music.",
"width": 1344,
"height": 768,
"duration": 11
}Response
[
{
"taskType": "videoInference",
"taskUUID": "9b2e5f71-c8a4-4d13-b6f0-2a7e1d94c358",
"videoUUID": "d6a1c3f8-7e25-4b90-a4d2-5c8f0e3b7a19",
"videoURL": "https://vm.runware.ai/video/os/a14d18/ws/2/vi/d6a1c3f8-7e25-4b90-a4d2-5c8f0e3b7a19.mp4"
}
]The 22-word line runs about nine seconds, so the request asks for 11 and leaves a beat on either side. Fitting the line to the duration covers the arithmetic.
What happens to sound you leave out
The model follows the brief literally, and that includes the soundtrack. Without an audio clause it still produces sound, but every choice about it is the model's own, and a music bed is one of the choices it can make. The pair below is one coworking scene, first with no audio direction at all, then with the sounds named.
Shot on a full-frame mirrorless camera with a 35mm lens at f/2.8, locked off on a tripod at desk height. A woman in her thirties in a rust-colored sweater sits at a long shared desk in a bright coworking space, closes her laptop, stands up and slings a canvas tote bag over her shoulder. Other members work at desks in the soft-focus background. Large windows, pale wood and green plants. No logos, brand names or printed words anywhere in frame.
Shot on a full-frame mirrorless camera with a 35mm lens at f/2.8, locked off on a tripod at desk height. A woman in her thirties in a rust-colored sweater sits at a long shared desk in a bright coworking space, closes her laptop, stands up and slings a canvas tote bag over her shoulder. Other members work at desks in the soft-focus background. Large windows, pale wood and green plants. No logos, brand names or printed words anywhere in frame. Audio: the soft click of the laptop lid closing, her chair rolling back across the concrete floor, low murmured conversation across the room, and a printer running somewhere behind her. No music, no dialogue.
The directed take has the laptop lid and the chair rolling back, each landing where the picture puts it, over the murmur of the room. The undirected take is a guess at what a coworking space sounds like. Rule out music unless you want a bed: end the prompt with No music, and add no dialogue whenever a person is in frame but should not speak.
Writing dialogue
Quoting the line
Write the spoken words inside quotation marks, exactly as they should be said, and attach them to a described person. The model speaks what is quoted and does not improvise around it.
Three things around the quote shape the performance:
- Who speaks. Tie the line to a person the prompt has already described, so there is no doubt whose mouth moves.
- How they say it. A short delivery note before the colon ("in a calm, clear, encouraging voice") sets tone and pace.
- What they do while speaking. The garden center owner lifts the plant toward the camera as she talks, and the action and the words land together.
Adding "her lips stay in sync with every word" after the quote is a cheap way to reinforce the sync on close framings, where a mismatch is most visible.
Fitting the line to the duration
Dialogue plays at roughly 2.5 words per second. That gives a word budget for every duration:
duration | Words that fit |
|---|---|
| 5 | about 12 |
| 8 | about 20 |
| 10 | about 25 |
| 12 | about 30 |
| 15 | about 37 |
Budget a little under the number so the speaker has a beat before and after the line. The pair below is the same 30-word skincare review, first squeezed into 5 seconds, then given 12.
Shot on a smartphone front camera held at arm's length, vertical, with a natural handheld wobble. A woman in her late twenties with freckles and her hair tied up in a claw clip sits on her bed in a sunlit bedroom, holding a small amber dropper bottle up next to her face, and talks to the camera like she is telling a friend: "Okay, so I've been using this vitamin C serum every morning for three weeks, and the dark spots on my cheeks are genuinely lighter. Two drops, pat it in, done." Her lips stay in sync with every word. Real skin with visible freckles and texture, no retouching, no beauty filter. No logos, brand names or printed words anywhere in frame, including on the bottle. Audio: her voice close to the phone microphone, a quiet bedroom room tone, and birdsong faintly outside the window. No music.
Shot on a smartphone front camera held at arm's length, vertical, with a natural handheld wobble. A woman in her late twenties with freckles and her hair tied up in a claw clip sits on her bed in a sunlit bedroom, holding a small amber dropper bottle up next to her face, and talks to the camera like she is telling a friend: "Okay, so I've been using this vitamin C serum every morning for three weeks, and the dark spots on my cheeks are genuinely lighter. Two drops, pat it in, done." Her lips stay in sync with every word. Real skin with visible freckles and texture, no retouching, no beauty filter. No logos, brand names or printed words anywhere in frame, including on the bottle. Audio: her voice close to the phone microphone, a quiet bedroom room tone, and birdsong faintly outside the window. No music.
In five seconds the model crams all thirty words in, at about six words a second, more than twice the pace of natural speech. At twelve seconds the same review plays at the pace a creator actually talks. Count the words before you pick the duration, so the pace is set by you and not by the clip length.
Two speakers
Describe each person so they can be told apart, and give each their own quoted line and delivery note. Then state that each person's lips move only on their own line.
Shot on a full-frame cinema camera with a 28mm lens at T4, locked off on a tripod at chest height. An empty, freshly painted main bedroom in an apartment, with pale oak floors and a wall of full-height windows on the left letting in morning sun. A real-estate agent in her forties in a navy blazer stands by the windows, facing a young couple in their late twenties who stand near the doorway. The agent gestures toward the window and says, warmly: "And this is the main bedroom. Morning light, and the closet runs the full length of that wall." The woman in the couple, in a green sweater, looks at the wall, then turns to her partner and says, delighted: "Oh, that's so much space." Each person's lips move only on their own line. No logos, brand names or printed words anywhere in frame. Audio: both voices natural and close, with the slight echo of an empty room, and faint traffic outside. No music.
The agent speaks first and the buyer answers, in the order the prompt gives them. The empty room also shows up in the sound: both voices carry the slight echo the audio clause asked for. Keep the total across both speakers inside the word budget, since the lines share one clip.
Other languages
A line in another language is written in that language inside the quotes, with the language named in the delivery note. The pair below is one pharmacy scene, localized.
Shot on a large-format cinema camera with a 50mm lens at T2.8, locked off on a tripod at eye level. A pharmacist in her fifties with short dark hair and reading glasses, in a white pharmacy coat, stands behind a pharmacy counter with neat shelves of plain white boxes behind her. She looks at the camera and says, kindly and clearly, in English: "Take one tablet every eight hours, always with food. If you feel dizzy, stop taking it and come back to see me." Her lips stay in sync with every word. Bright, even overhead light. Real skin texture, no retouching. No logos, brand names or printed words anywhere in frame, including on the boxes and her coat. Audio: her voice close and clear, a quiet pharmacy room tone, and the soft chime of the shop door far behind the camera. No music.
Shot on a large-format cinema camera with a 50mm lens at T2.8, locked off on a tripod at eye level. A pharmacist in her fifties with short dark hair and reading glasses, in a white pharmacy coat, stands behind a pharmacy counter with neat shelves of plain white boxes behind her. She looks at the camera and says, kindly and clearly, in Spanish from Spain: "Tómese una pastilla cada ocho horas, siempre con comida. Si nota mareos, deje de tomarla y vuelva a verme." Her lips stay in sync with every word. Bright, even overhead light. Real skin texture, no retouching. No logos, brand names or printed words anywhere in frame, including on the boxes and her coat. Audio: her voice close and clear, a quiet pharmacy room tone, and the soft chime of the shop door far behind the camera. No music.
Only the quoted line and the language note change between the two prompts. Keep everything else identical when you localize, so the versions differ in language and not in staging.
For a presenter who has to be the same person in every language, give the model a reference image, covered in the reference guide.
Placing sound by distance
Ambience reads as real when it has depth. Name each sound and where it sits: close to the camera, across the room, far away. The warehouse prompt below lists its audio by distance, nearest first.
Shot on a large-format cinema camera with a 35mm lens at T4, locked off on a tripod at eye level, looking down a long aisle of tall steel pallet racking in a large distribution warehouse. A warehouse worker in a hi-vis orange vest and a white hard hat stands in the foreground on the right, holding a tablet at his side and looking down the aisle. Far down the aisle, a forklift crosses from left to right carrying a wrapped pallet. Cool, even overhead LED light. No logos, brand names or printed words anywhere in frame, including on the vest, the racking and the forklift. Audio, close to the camera: the low hum of the building's ventilation. At mid distance: a pallet jack rolling over a floor seam. Far away: the forklift's reversing beeper and the whine of its electric motor as it crosses the aisle. No music, no dialogue.
The forklift is small in the frame, and its beeper is quiet and distant to match. That agreement between picture and sound is what the shared pass buys. Writing the distances down is what makes it deliberate.
Timing an effect to the action
An effect lands on cue when the prompt ties it to the moment in the action that causes it: "the clink the moment the pour stops", "the latch as the door closes". Describing a sound on its own gets you the sound somewhere in the clip.
Shot on a large-format cinema camera with a 100mm macro lens at T2.8, locked off on a tripod at counter height. A tall clear glass with ice and a slice of lime stands on a pale travertine counter in soft morning light. A hand enters from the right holding a small glass carafe and pours a short stream of sparkling water into the glass, then lifts the carafe away. Bubbles rise and fizz in the glass. No logos, brand names or printed words anywhere in frame. Audio: the pour splashing onto the ice, a fizz of bubbles that grows louder as the pour ends, and the ice settling with a single clear clink the moment the pour stops. A quiet kitchen room tone underneath. No music, no dialogue.
The audio clause follows the order of the action: the splash during the pour, the fizz as it ends, the clink when it stops. Writing the sounds in sequence gives the model a timeline to follow.
Music beds
When a clip should carry music, describe the track the way a music supervisor would brief it: genre, tempo, the instruments and where it sits in the mix. A vague "upbeat music" leaves all of that to the model.
Shot on a large-format cinema camera with a 75mm lens at T2, locked off on a tripod at waist height. A model in her twenties in an oversized cream trench coat, wide black trousers and white leather sneakers walks toward the camera across an empty concrete rooftop at golden hour, the city skyline soft behind her. Her hair is pulled back into a sleek low bun. She stops a few steps from the camera, turns her head to the left, then back. Warm low sun from behind, a clean, muted color grade. No logos, brand names or printed words anywhere in frame. Audio: a minimal house track at 120 beats per minute with a deep, soft kick drum, crisp hi-hats and a warm synth chord on every other bar, mixed forward, with her footsteps on the concrete faintly underneath. No dialogue.
"Mixed forward, with her footsteps faintly underneath" sets the balance between the music and the scene. For a clip that will be cut to your own licensed track later, skip the bed entirely: ask for No music and let the footsteps and room carry it.
To use a track you already have, rather than one described in words, pass it as an audio reference. The audio references guide covers using your own music or voiceover as the clip's soundtrack.
Tips
-
Always direct the audio. End the prompt with an
Audio:clause. Without one, every sound is the model's choice, music included. -
Exclude music and speech explicitly. Write
No musicandno dialoguewhen you mean them. The model does not add speech you did not write, but naming the exclusion removes any doubt on a shot with a person in it. -
Quote the line exactly and attach it to a described person. Add a short delivery note before the quote to set tone and pace.
-
Budget about 2.5 words per second. Count the words first, then pick the duration, and leave a beat at each end.
-
Give each speaker their own line and description. State that each person's lips move only on their own line.
-
Write non-English lines in the target language. Name the language in the delivery note and keep the rest of the prompt identical across versions.
-
Place ambience by distance. Close, mid and far, nearest first. Sounds that match the size of their source in the frame are what make a scene feel real.
-
Tie each effect to its moment in the action. "The clink the moment the pour stops" lands on cue. "A clink" lands anywhere.