P-Video-2

P-Video-2 is Pruna AI's quality-focused successor to P-Video, creating clips from text, images, or audio-conditioned inputs through one model. It produces sharper close-ups and foreground subjects, stronger identity and input-image consistency, improved lip synchronization for natively generated speech, and clearer on-screen text while generating dialogue, music, and sound effects with the video. It supports optional first-and-last-frame guidance, draft iteration, durations up to 20 seconds, 720p or 1080p output, and 24 or 48 FPS.

Complete technical specification for integration
Ready-to-use code snippets for common workflows
Step-by-step tutorials for advanced use cases
← All GuidesWriting prompts for scene, motion and sound
How to write prompts for Pruna P-Video-2, from the clause order the model reads to the framing, camera motion and sound directions that ship inside the same string.
Introduction
One string does the whole job here. positivePrompt carries the subject, the framing, the camera, the light and the sound that comes back inside the file, because P-Video-2 generates its own audio track unless you switch it off. A prompt written for a silent model leaves that last decision to the model.
The prompt is capped at 2048 characters, which is roughly 300 words. That is generous enough to over-write, and the failure mode of a long prompt is not truncation but dilution, where the model spreads its attention across instructions you did not need.
A two-shot from the bonnet of a parked car looking in through the windscreen at two people in the front seats at night, both facing forward. The driver says, "We can still turn around." The passenger, without looking at him, answers, "No we cannot." Street light falling across the glass, rain beaded on the windscreen. Photoreal television drama cinematography, locked-off camera, no text, no logos. Audio: their two natural speaking voices taking turns, quiet interior car tone underneath, no music.
Turn the sound on. Both lines are in the prompt, word for word, and they come back spoken in turn by two different voices with the mouths matching. The picture and the dialogue were generated together from one string, so nothing had to be synced afterwards. This guide covers the clause order the model reads, how framing changes the detail you get back, how much camera movement holds, and the production patterns worth copying.
What the model reads
A P-Video-2 prompt is a stack of layers rather than a sentence. The model reads all of it, but the layers do different jobs, and leaving one out hands that decision to the prompt expander described at the end of this guide, which is on by default. The hero prompt above, broken into its parts:
Two of those layers are worth calling out because they are the ones people leave off.
The exclusion layer earns its place on anything commercial. A model that has learned from stock footage will add a logo, a caption or a watermark to a product shot given the chance, and "no text anywhere in frame" is cheaper than discovering it after the render. The same clause handles the people you did not ask for, which is the usual surprise in an interior shot.
The audio layer is not optional in practice, only in syntax. settings.audio defaults to true, so a prompt with nothing to say about sound still comes back with sound, invented to match the picture. Writing Audio: and then describing what you want is how you stop that being a surprise. Native audio covers what the clause can do once you are deliberately using it.
Order matters less than presence, but the shape that reads most reliably is scene first, camera second, sound last. Establish what is in front of the lens, then say how it is being filmed, then say what it sounds like.
Framing decides how much detail you get
P-Video-2 is built around close subjects, and the difference is not subtle. The same subject at three framings, on one seed so nothing else moves:
A tight chest-up shot of a woman in her thirties wearing a camel wool overcoat over a cream rollneck, turning her shoulders slowly toward camera as the coat lapel catches the light. Bright concrete gallery interior behind her, thrown well out of focus. Soft overcast daylight from a tall window on the left. Photoreal fashion editorial cinematography, slow subject movement, locked-off camera. Audio: quiet interior room tone, soft fabric movement, no music, no voice.
A waist-up shot of a woman in her thirties wearing a camel wool overcoat over a cream rollneck, turning her shoulders slowly toward camera as the coat lapel catches the light. Bright concrete gallery interior behind her, softly out of focus. Soft overcast daylight from a tall window on the left. Photoreal fashion editorial cinematography, slow subject movement, locked-off camera. Audio: quiet interior room tone, soft fabric movement, no music, no voice.
A wide full-length shot of a woman in her thirties wearing a camel wool overcoat over a cream rollneck, walking slowly toward camera down a bright concrete gallery corridor, the coat moving with her stride. Soft overcast daylight from a run of tall windows on the left. Photoreal fashion editorial cinematography, subject walking, locked-off camera. Audio: quiet interior room tone, soft footsteps on concrete, no music, no voice.
Watch the lapel across the three. Close in, the weave and the stitch line are rendered. At waist height the garment keeps its shape and the surface goes generic. Full-length, the coat is a silhouette that moves correctly and carries almost no material information at all.
That has a direct consequence for anyone shooting product or garment work: the shot that sells the material has to be a close one. If the deliverable needs both, generate them as separate clips at their own framings rather than asking one wide shot to carry detail it cannot resolve.
Foreground objects get the same benefit as close subjects. A prompt that puts the hero item near the lens with the scene falling away behind it holds detail that the same item, described as part of the room, will lose.
How much camera movement holds
Subject motion and camera motion are separate instructions and the model treats them differently. Three camera directions over one interior, same seed:
A locked-off wide shot of a bright modern living room with a pale grey linen sofa, a low walnut coffee table on a wool rug, and tall windows with sheer curtains moving very slightly in the air. Warm late afternoon daylight across a pale oak floor. Photoreal real-estate interior cinematography, tripod-mounted camera that does not move, no people. Audio: quiet interior room tone, faint movement of fabric, no music, no voice.
A slow steady dolly moving forward through a bright modern living room, past a pale grey linen sofa and a low walnut coffee table on a wool rug toward tall windows with sheer curtains. Warm late afternoon daylight across a pale oak floor. Photoreal real-estate interior cinematography, smooth continuous forward camera movement at walking pace, no people. Audio: quiet interior room tone, faint movement of fabric, no music, no voice.
A fast whip pan across a bright modern living room that snaps to a halt on a pale grey linen sofa, then immediately cranes up and over the coffee table while rotating a full ninety degrees, then drops back down to floor level and pushes hard through the doorway into the hall. Warm late afternoon daylight, pale oak floor. Photoreal real-estate interior cinematography, aggressive handheld camera, no people. Audio: quiet interior room tone, no music, no voice.
The locked-off and the dolly both come back clean. The third asks for four camera moves inside five seconds and the room stops being a room, because the model is inventing geometry faster than it can keep it consistent. Pruna are explicit that extreme cinematic camera motion is outside what this model was built for, and the practical reading of that is a budget: one camera move per clip, described once.
The budget is spent per clip, not per second, so a longer duration does not buy you a second move. If a shot genuinely needs to travel and then turn, that is two clips and a cut in your edit.
Subject motion is the cheaper way to get life into a frame. A locked-off camera with something moving inside it, a curtain, a hand, a walk across frame, is the most reliable shape this model produces, and it is also what most social and listing formats actually want.
Lighting is a prompt layer, not a post step
Light is one of the few descriptions that changes the whole render for almost no prompt length. The same testimonial setup with one clause swapped:
A chest-up shot of a man in his forties with a short greying beard, wearing a charcoal crew-neck jumper, sitting at a plain oak table and listening, nodding once slowly. Large soft window light from camera left wrapping around his face, gentle fall-off, open shadows, plain pale wall behind. Photoreal corporate testimonial cinematography, locked-off camera, shallow depth of field. Audio: quiet room tone, no music, no voice.
A chest-up shot of a man in his forties with a short greying beard, wearing a charcoal crew-neck jumper, sitting at a plain oak table and listening, nodding once slowly. Strong low afternoon sun coming through a window blind to his left, cutting bright hard stripes across his face and the wall behind, with sharp dark bands between them. Photoreal corporate testimonial cinematography, locked-off camera, shallow depth of field. Audio: quiet room tone, no music, no voice.
Both are single clauses and they produce two different pieces of brand work.
What the model responds to is the effect you can see, not the technique that produces it. "Hard stripes across his face with sharp dark bands between them" lands. The lighting-department names for the same thing do not: split lighting, Rembrandt and chiaroscuro all came back as an ordinary dim room. Describe the shadow, the direction it falls and how hard its edge is, and leave the name of the setup out of it.
The model will happily put the lamp in shot. Asking for a hard key from the left tends to produce a visible lamp in the corner of frame, and "the light source stays out of frame" does not reliably prevent it. Describing the light by what it does to the subject, rather than by the fixture making it, is the way around that.
Asking for cuts
A prompt is not limited to one continuous take. Name the cuts and the model will make them, which is worth knowing because most video models will not.
A product designer in a mustard cardigan sits at a standing desk in a bright open-plan office, looking at a large monitor showing a colourful dashboard, then turns her head to her colleague off-screen and smiles. One continuous unbroken shot, no cuts. Slow gentle push in on the desk. Soft daylight from a window bank behind her, plants along the sill. Photoreal SaaS brand-film cinematography. Audio: quiet office room tone, faint keyboard, no music, no voice.
Three cuts. First shot: a product designer in a mustard cardigan at a standing desk looking at a dashboard on a large monitor. Cut to a second shot: a close-up of her hand moving a mouse. Cut to a third shot: a wide of the whole open-plan office with six people at desks. Bright daylight throughout. Photoreal SaaS brand-film cinematography. Audio: quiet office room tone, faint keyboard, no music, no voice.
The three-cut prompt returns three shots: the designer at her desk, a close-up of her hand on the mouse, then a wide of the office. The cuts are clean, and the designer survives them, which is the part that usually fails elsewhere.
Two things make it work. Number the shots ("First shot… Cut to a second shot… Cut to a third shot") rather than describing a sequence of events and hoping. And give each shot its own framing, because a cut between two similar framings reads as a jump rather than an edit.
Where this stops being the right tool is length. Three cuts inside six seconds gives each shot two seconds, so it suits a title sequence or a social bumper rather than a narrative beat. When each shot needs room to breathe, generate them separately and cut them yourself, and use image-to-video so every clip starts from the same still.
The reverse is the more common need. "One continuous unbroken shot, no cuts" is worth writing explicitly whenever you want a single take, because a prompt that describes two things happening in sequence can be read as an invitation to cut.
Production patterns
The two shapes below cover most of what gets shipped from a prompt-only workflow. Both are complete requests.
Vertical social with a caption safe area. Vertical formats get a caption in post almost every time, so the prompt has to reserve the space. Naming the headroom in the framing clause keeps the subject clear of the overlay instead of fighting it.
import { createClient } from '@runware/sdk'
const client = await createClient({ apiKey: process.env.RUNWARE_API_KEY })
await client.connect()
const [result] = await client.run({
model: 'prunaai:p-video@2',
positivePrompt: 'A woman in charcoal leggings and a coral crop top holds a forearm plank on a dark grey mat in a bright home studio, her body in one straight line, shoulders steady, a slow controlled breath moving her ribs. Vertical framing for social, subject centred with headroom for a caption. Low side daylight from a window on the right raking across the mat, plain white wall behind. Photoreal fitness-app cinematography, locked-off camera at mat height. Audio: calm room tone, steady controlled breathing, no music, no voice.',
width: 704,
height: 1280,
duration: 6,
deliveryMethod: 'async'
})import asyncio
import os
from runware import Runware
async def main():
async with Runware(api_key=os.environ["RUNWARE_API_KEY"]) as client:
results = await client.run({
"model": "prunaai:p-video@2",
"positivePrompt": "A woman in charcoal leggings and a coral crop top holds a forearm plank on a dark grey mat in a bright home studio, her body in one straight line, shoulders steady, a slow controlled breath moving her ribs. Vertical framing for social, subject centred with headroom for a caption. Low side daylight from a window on the right raking across the mat, plain white wall behind. Photoreal fitness-app cinematography, locked-off camera at mat height. Audio: calm room tone, steady controlled breathing, no music, no voice.",
"width": 704,
"height": 1280,
"duration": 6,
"deliveryMethod": "async"
})
asyncio.run(main())curl https://api.runware.ai/v1 \
-H "Authorization: Bearer $RUNWARE_API_KEY" \
-H "Content-Type: application/json" \
-d '[
{
"taskType": "videoInference",
"taskUUID": "c74e1b93-2a68-4f05-b3d1-6e9a5c2f8047",
"model": "prunaai:p-video@2",
"positivePrompt": "A woman in charcoal leggings and a coral crop top holds a forearm plank on a dark grey mat in a bright home studio, her body in one straight line, shoulders steady, a slow controlled breath moving her ribs. Vertical framing for social, subject centred with headroom for a caption. Low side daylight from a window on the right raking across the mat, plain white wall behind. Photoreal fitness-app cinematography, locked-off camera at mat height. Audio: calm room tone, steady controlled breathing, no music, no voice.",
"width": 704,
"height": 1280,
"duration": 6,
"deliveryMethod": "async"
}
]'runware run prunaai:p-video@2 \
positivePrompt="A woman in charcoal leggings and a coral crop top holds a forearm plank on a dark grey mat in a bright home studio, her body in one straight line, shoulders steady, a slow controlled breath moving her ribs. Vertical framing for social, subject centred with headroom for a caption. Low side daylight from a window on the right raking across the mat, plain white wall behind. Photoreal fitness-app cinematography, locked-off camera at mat height. Audio: calm room tone, steady controlled breathing, no music, no voice." \
width=704 \
height=1280 \
duration=6 \
deliveryMethod=async{
"taskType": "videoInference",
"taskUUID": "c74e1b93-2a68-4f05-b3d1-6e9a5c2f8047",
"model": "prunaai:p-video@2",
"positivePrompt": "A woman in charcoal leggings and a coral crop top holds a forearm plank on a dark grey mat in a bright home studio, her body in one straight line, shoulders steady, a slow controlled breath moving her ribs. Vertical framing for social, subject centred with headroom for a caption. Low side daylight from a window on the right raking across the mat, plain white wall behind. Photoreal fitness-app cinematography, locked-off camera at mat height. Audio: calm room tone, steady controlled breathing, no music, no voice.",
"width": 704,
"height": 1280,
"duration": 6,
"deliveryMethod": "async"
}Response
[
{
"taskType": "videoInference",
"taskUUID": "c74e1b93-2a68-4f05-b3d1-6e9a5c2f8047",
"videoUUID": "5b8d3f21-9c47-4a60-8e15-2d7a9c4f1b83",
"videoURL": "https://vm.runware.ai/video/os/a14d18/ws/2/vi/5b8d3f21-9c47-4a60-8e15-2d7a9c4f1b83.mp4"
}
]A woman in charcoal leggings and a coral crop top holds a forearm plank on a dark grey mat in a bright home studio, her body in one straight line, shoulders steady, a slow controlled breath moving her ribs. Vertical framing for social, subject centred with headroom for a caption. Low side daylight from a window on the right raking across the mat, plain white wall behind. Photoreal fitness-app cinematography, locked-off camera at mat height. Audio: calm room tone, steady controlled breathing, no music, no voice.
Flat motion design. Photoreal is the default the model falls back on, so an illustrated deliverable has to say so in the style clause and then keep saying so. Naming the palette and ruling out gradients is what stops a flat vector brief drifting into rendered 3D.
A flat vector animation of one single cardboard parcel box travelling slowly from left to right along a plain horizontal conveyor belt, with a progress bar filling underneath it. Limited palette of navy, coral and off-white, clean geometric shapes, flat colour with no gradients and no texture. Motion-design brand animation, even flat lighting, locked-off camera, exactly one box and nothing else on the belt, no text. Audio: light rhythmic ticking that matches the belt, one soft chime as the progress bar fills, no music, no voice.
The sound direction in that one is doing something the picture cannot. Ticking that matches the belt and a chime on the fill gives the animation its timing, and it costs one clause.
The prompt expander
settings.promptUpsampling defaults to true, so unless you turn it off your text is rewritten before the model acts on it, and every layer you left blank is filled in for you. It picks the setting, the light, the lens, the camera move and the sound, and it picks differently on every call.
{
"settings": {
"promptUpsampling": false
}
}Every clip below uses the same four-word prompt, "Show the running shoes", on the same pinned seed. Only the flag changes:
Show the running shoes.
Show the running shoes.
Four words name a subject and nothing else. With the expander off, that is exactly what you get: a pair of shoes, worn, in ordinary daylight, and no production around them. With it on, the rewrite supplies a studio floor, a polished surface and campaign lighting. The clip is better looking, and not one of those choices was yours.
Now two more calls, still with the expander on and still on that same seed:
Show the running shoes.
Show the running shoes.
A pinned seed is what makes a run repeatable, and here it does nothing: four calls, one seed, four different shoots. The rewrite happens before the seed is used, so the seed is applied to a different prompt every time.
One more thing worth knowing: the expanded prompt is not returned in the response, so when a result surprises you there is nothing to inspect.
That gives a simple rule. Leave it on while you are exploring and have not written a brief yet. Turn it off for anything branded, anything batched, and anything you are debugging, where an invented set or a different light on every call is the problem rather than the feature. A full prompt with the flag off is what makes forty clips look like one shoot.
Tips
-
Write the audio clause even when you want silence. Sound is generated by default, so "no music, no voice" is an instruction, not a comment. Leaving it out is how an explainer arrives with a stock music bed under it.
-
Spend the camera budget once. One move per clip. A prompt with a pan and a crane and a push produces geometry that does not survive the move.
-
Put the hero close to the lens. Material, text on packaging and facial detail all resolve at close framings and dissolve at full-length ones.
-
Say whether you want cuts. "One continuous unbroken shot, no cuts" holds a single take. Numbered shots ("First shot… Cut to a second shot…") gets you an edit. Leave it unsaid and the model picks.
-
Rule out what you do not want. No text, no logos, no people, no captions. Commercial footage in the training data is full of all four.
-
Describe light by size and direction. "Large soft key from camera left, no fill" changes more about a shot than any adjective, and it reads the same way to the model as it does to a crew.
-
Pin a seed while you iterate on wording. Changing one clause against a fixed seed shows you what that clause did. Changing wording and seed together shows you nothing. Draft mode covers the loop this belongs to.
-
Turn the expander off for branded and batched work.
settings.promptUpsamplingdefaults totrueand rewrites your prompt differently every call. A full prompt with it off is what keeps a set consistent. -
Keep prompts under about 200 words. The cap is 2048 characters, and long past that point extra clauses compete with each other rather than adding up.