P-Video-2-Pro

P-Video-2-Pro is Pruna AI's quality-tier video generation model built on MiniMax H3, creating clips from text or from a first frame with an optional last frame. It handles multi-beat camera moves, heavy physics like water, fire, and fabric, coherent full-body motion, and two-shot dialogue with lip sync, generating audio with every clip. It offers a speed or quality recipe, three levels of prompt expansion, durations from 5 to 15 seconds, and 480p or 768p output at 24 FPS.

Complete technical specification for integration
Ready-to-use code snippets for common workflows
Step-by-step tutorials for advanced use cases
← All GuidesSpeed, quality, size and length
How to pick speed or quality mode in Pruna P-Video-2-Pro, choose between 480p and 768p, and set a clip length between 5 and 15 seconds.
Introduction
Three settings decide what a P-Video-2-Pro clip costs before the prompt says anything: settings.mode, the resolution tier and duration. They multiply per second of output, and the cheapest combination runs at roughly a quarter of the rate of the most expensive one.
A surfer in a black wetsuit drops down the face of a head-high turquoise wave and carves a hard turn at the bottom, throwing a long fan of spray off the tail of the board as the lip curls over behind her. The camera tracks alongside from the water at board height. Bright late morning sun, spray catching the light. Photoreal surf brand cinematography, no text, no logos. Audio: the roar of the breaking wave, the hiss of the board carving through water, no music, no voice.
A surfer in a black wetsuit drops down the face of a head-high turquoise wave and carves a hard turn at the bottom, throwing a long fan of spray off the tail of the board as the lip curls over behind her. The camera tracks alongside from the water at board height. Bright late morning sun, spray catching the light. Photoreal surf brand cinematography, no text, no logos. Audio: the roar of the breaking wave, the hiss of the board carving through water, no music, no voice.
One prompt, one seed and the expander off, rendered in both modes. The mode is the only difference between them, and it is enough to re-stage the shot as well as roughly double the per-second rate. This guide covers what quality mode buys, the two resolution tiers, what extra length does, what a seed carries between settings, and a loop for exploring cheaply and delivering once.
Speed or quality
settings.mode defaults to speed, the faster and cheaper recipe. quality is slower, and it bills at roughly twice the per-second rate of speed at either tier.
{
"settings": {
"mode": "quality"
}
}Fine detail in motion is where the two separate. The same haircare shot in both modes, on one seed with the expander off:
A close-up of a woman with long dark curly hair under a single studio light as she tosses her head to one side, the curls swinging out and bouncing back into place, individual strands catching the light against a plain black background. Photoreal haircare advertising cinematography, locked-off camera, no text, no logos. Audio: a soft rush of hair movement, quiet studio room tone, no music, no voice.
A close-up of a woman with long dark curly hair under a single studio light as she tosses her head to one side, the curls swinging out and bouncing back into place, individual strands catching the light against a plain black background. Photoreal haircare advertising cinematography, locked-off camera, no text, no logos. Audio: a soft rush of hair movement, quiet studio room tone, no music, no voice.
Hair is a hard test because every curl is fine detail that moves. The two clips are separate takes, as the seed section below explains, so compare the texture rather than the frame. Quality renders the curls as clean, defined spirals that keep their shape as they swing. Speed gets the same toss across with a busier, grainier texture that looks crunchy up close.
Speed answers whether the prompt works: whether the move and the timing land. Quality is for the finals, where the surface of the subject is what the viewer is looking at.
Speed is also the faster one to wait for. Pruna's own figures put a 768p speed render at under a second of processing per second of video, and a quality render at a little over twice that.
480p or 768p
State the tier either with resolution, which takes 480p or 768p and defaults to 768p, or with an exact width and height pair. The two are mutually exclusive, and a request carrying both is rejected.
resolution on its own returns a 16:9 clip, so a different shape in text-to-video needs the pair. With a first frame in the request it is the other way around: the pair is rejected and the still sets the shape, as starting and ending on frames you choose covers.
The same product shot at both tiers, on one seed with the expander off:
An overhead close-up of an open glass jar of whipped white body butter on a pale pink surface. A fingertip swipes slowly through the surface, lifting a soft glossy peak that holds its shape. Soft diffused studio light. Photoreal skincare e-commerce cinematography, square framing, locked-off camera, no text, no logos. Audio: a faint soft swipe, quiet room tone, no music, no voice.
An overhead close-up of an open glass jar of whipped white body butter on a pale pink surface. A fingertip swipes slowly through the surface, lifting a soft glossy peak that holds its shape. Soft diffused studio light. Photoreal skincare e-commerce cinematography, square framing, locked-off camera, no text, no logos. Audio: a faint soft swipe, quiet room tone, no music, no voice.
The square is used here because it is the same shape at both tiers. With the expander off the seed also holds the staging across tiers, so both clips follow the same swipe and the difference is detail. At 768p the glossy edge of the peak and the whipped texture around it resolve, and on a product page that texture is what is being sold.
Ship at 768p, the tier for anything a customer sees full size and the one where surface detail survives. 480p is for exploring, and for placements that will only ever play small.
The tier name is the short edge of every size, so a 768p square is 768 × 768 and carries fewer pixels than a 1344 × 768 landscape frame. Shapes also shift slightly between tiers: only 4:3, 3:4 and 1:1 keep their exact shape from 480p to 768p, and a 16:9 clip explored at 832 × 480 comes back a little wider at 1344 × 768, so framing that sat tight to an edge in the draft can move in the final.
None of the 16:9 sizes is exactly 16:9. A 1344 × 768 clip dropped into a 1920 × 1080 sequence needs a scale and a thin crop, so plan that pass for any deliverable with a fixed frame.
Clip length
duration is a whole number of seconds from 5 to 15, and it defaults to 5. Every second is billed, so length multiplies with the mode and the tier. The returned clip lands near the requested length rather than on it: 5 seconds came back as 5.17, and 15 as 14.38. The same single-action prompt at both ends of the range:
A golden retriever sprints across a sunny backyard lawn and leaps to catch a red flying disc in its mouth, then lands and trots back toward the camera with it. Photoreal pet brand advertising cinematography, low tracking camera, warm afternoon light, no text, no logos, no people visible. Audio: paws thudding on grass, a happy bark after the catch, birdsong, no music, no voice.
A golden retriever sprints across a sunny backyard lawn and leaps to catch a red flying disc in its mouth, then lands and trots back toward the camera with it. Photoreal pet brand advertising cinematography, low tracking camera, warm afternoon light, no text, no logos, no people visible. Audio: paws thudding on grass, a happy bark after the catch, birdsong, no music, no voice.
Five seconds fits the prompt: a sprint, a catch and a trot back. At fifteen, the model stretches the same three actions rather than adding any. The catch lands at around six seconds and the trot back fills most of the second half, which suits a slow lifestyle clip and wastes seconds in anything cut to a script.
Write as much action as the length holds. Five seconds fits one action. Fifteen wants a sequence, or several camera beats like the ones in camera moves, physics and full-body motion.
What a seed carries
seed takes an integer, and it only makes a run repeatable while nothing else rewrites the request. The expander rewrites the prompt differently on every call, so turn it off before comparing settings on a pinned seed. With it off, the same request comes back as a byte-identical file, and what survives a change of setting depends on which setting changes.
Across tiers, the staging holds. The body butter swipe at 480p and 768p follows the same path at the same moments, and only the detail changes.
Across modes, the seed starts a new take. The surf clip at the top of this guide keeps the brief in both modes, a drop, a carve and a fan of spray at about the same moment, but the wave and the framing differ. Speed and quality are different recipes, and one seed means something different to each.
That sets the order of work: pick the mode before you pick the take. When a speed take is the one you want, lock its frame as a still and render the final from it, as starting and ending on frames you choose covers.
Exploring cheaply, delivering once
Per second of output, relative to 480p in speed mode:
| Mode | 480p | 768p |
|---|---|---|
speed | 1× | about 1.75× |
quality | 2× | about 3.75× |
A final at 768p in quality costs close to four times a 480p speed draft of the same length, and that gap is the budget for exploring. Spend it in the order the seed allows, with the expander off throughout:
- Explore at 480p in speed until the prompt, the duration and the timing are right.
- Switch to quality at 480p and try seeds until one gives you the take you want.
- Render that seed at 768p in quality, and the staging carries across.
import { createClient } from '@runware/sdk'
const client = await createClient({ apiKey: process.env.RUNWARE_API_KEY })
await client.connect()
const [result] = await client.run({
model: 'prunaai:p-video@2-pro',
positivePrompt: 'A surfer in a black wetsuit drops down the face of a head-high turquoise wave and carves a hard turn at the bottom, throwing a long fan of spray off the tail of the board as the lip curls over behind her. The camera tracks alongside from the water at board height. Bright late morning sun, spray catching the light. Photoreal surf brand cinematography, no text, no logos. Audio: the roar of the breaking wave, the hiss of the board carving through water, no music, no voice.',
width: 1344,
height: 768,
duration: 6,
seed: 7340216,
settings: {
mode: 'quality',
promptUpsampling: 'off'
},
deliveryMethod: 'async'
})import asyncio
import os
from runware import Runware
async def main():
async with Runware(api_key=os.environ["RUNWARE_API_KEY"]) as client:
results = await client.run({
"model": "prunaai:p-video@2-pro",
"positivePrompt": "A surfer in a black wetsuit drops down the face of a head-high turquoise wave and carves a hard turn at the bottom, throwing a long fan of spray off the tail of the board as the lip curls over behind her. The camera tracks alongside from the water at board height. Bright late morning sun, spray catching the light. Photoreal surf brand cinematography, no text, no logos. Audio: the roar of the breaking wave, the hiss of the board carving through water, no music, no voice.",
"width": 1344,
"height": 768,
"duration": 6,
"seed": 7340216,
"settings": {
"mode": "quality",
"promptUpsampling": "off"
},
"deliveryMethod": "async"
})
asyncio.run(main())curl https://api.runware.ai/v1 \
-H "Authorization: Bearer $RUNWARE_API_KEY" \
-H "Content-Type: application/json" \
-d '[
{
"taskType": "videoInference",
"taskUUID": "2e8b5c41-7d96-4a03-b1f8-6c4d9a3e7b52",
"model": "prunaai:p-video@2-pro",
"positivePrompt": "A surfer in a black wetsuit drops down the face of a head-high turquoise wave and carves a hard turn at the bottom, throwing a long fan of spray off the tail of the board as the lip curls over behind her. The camera tracks alongside from the water at board height. Bright late morning sun, spray catching the light. Photoreal surf brand cinematography, no text, no logos. Audio: the roar of the breaking wave, the hiss of the board carving through water, no music, no voice.",
"width": 1344,
"height": 768,
"duration": 6,
"seed": 7340216,
"settings": {
"mode": "quality",
"promptUpsampling": "off"
},
"deliveryMethod": "async"
}
]'runware run prunaai:p-video@2-pro \
positivePrompt="A surfer in a black wetsuit drops down the face of a head-high turquoise wave and carves a hard turn at the bottom, throwing a long fan of spray off the tail of the board as the lip curls over behind her. The camera tracks alongside from the water at board height. Bright late morning sun, spray catching the light. Photoreal surf brand cinematography, no text, no logos. Audio: the roar of the breaking wave, the hiss of the board carving through water, no music, no voice." \
width=1344 \
height=768 \
duration=6 \
seed=7340216 \
settings.mode=quality \
settings.promptUpsampling=off \
deliveryMethod=async{
"taskType": "videoInference",
"taskUUID": "2e8b5c41-7d96-4a03-b1f8-6c4d9a3e7b52",
"model": "prunaai:p-video@2-pro",
"positivePrompt": "A surfer in a black wetsuit drops down the face of a head-high turquoise wave and carves a hard turn at the bottom, throwing a long fan of spray off the tail of the board as the lip curls over behind her. The camera tracks alongside from the water at board height. Bright late morning sun, spray catching the light. Photoreal surf brand cinematography, no text, no logos. Audio: the roar of the breaking wave, the hiss of the board carving through water, no music, no voice.",
"width": 1344,
"height": 768,
"duration": 6,
"seed": 7340216,
"settings": {
"mode": "quality",
"promptUpsampling": "off"
},
"deliveryMethod": "async"
}Response
[
{
"taskType": "videoInference",
"taskUUID": "2e8b5c41-7d96-4a03-b1f8-6c4d9a3e7b52",
"videoUUID": "d49a2f63-1b85-4e07-9c3a-8f6e2b5d1c79",
"videoURL": "https://vm.runware.ai/video/os/a14d18/ws/2/vi/d49a2f63-1b85-4e07-9c3a-8f6e2b5d1c79.mp4"
}
]Tips
-
Explore in speed, choose the take in quality. A seed does not carry a shot from one mode to the other, so the take you approve has to come from the mode you deliver in.
-
Ship at 768p. 480p is a drafting tier and a tier for small placements.
-
Pick one way to state the size. Send
resolutionor thewidthandheightpair, never both, and remember thatresolutionalone gives 16:9. -
Draft in the ratio you ship. Only 4:3, 3:4 and 1:1 keep their exact shape between 480p and 768p, so a 16:9 draft previews a slightly narrower frame.
-
Budget for the 16:9 mismatch.
1344 × 768is 1.75:1, so a fixed 1920 × 1080 frame needs a scale and a crop. -
Match the action to the length. Five seconds holds one action. A longer clip stretches the same actions instead of adding new ones.
-
Change the tier, not the mode, after approval. With the expander off, a seed holds its staging from 480p to 768p but not from speed to quality.
-
Count the multiplier before a batch. Mode, tier and seconds multiply, and a quality 768p final costs close to four times a speed 480p draft.