
Text to image
The foundational pipeline: generate images from text prompts using diffusion models, with control over every step of the process.
Understand core generative AI concepts and explore model-specific guides with practical examples and code.
Deep dives into every parameter and feature of the image generation pipeline. Each page covers how a feature works, its visual impact, and practical usage patterns with interactive examples.
Step-by-step tutorials built around specific models. Each guide walks through a real use case with code, prompts, and results you can reproduce.
How to compose FLUX 3's synchronized audio: layering ambient and effects, directing music by instrumentation and tempo, directing speech delivery, and rendering multilingual dialogue with clean lip-sync.

How to build multi-shot video sequences inside one FLUX 3 generation: HARD CUT syntax, shot contrast, time compression, threading an audio bed across the cuts, and pacing shot rhythm.

How to edit a finished clip with a prompt in MiniMax H3: replace a subject, relight a scene, add or remove elements, and combine edits, keeping the rest untouched.

How to transfer motion, a camera move, or an acting performance from a reference video onto a new subject with MiniMax H3 Omni Reference.

How to use reference images on Qwen Image 3.0 to recolour a packshot, carry a product into a new scene, or stage two products together in one shot.

How to design posters, packaging, ads, and menus with P-Image-Ideogram, and the workflow that gets on-image copy to render clean and exact.

How to write multi-shot storyboard prompts for HappyHorse 1.1 to direct cinematic sequences with shot sizes, camera movement, and subject continuity in a single call.
How to use HappyHorse 1.1's reference workflow to cast one or more characters into a generated video and preserve their identity through every cut.
How to generate infographics, comparison grids, newspaper pages, and interface mockups with Qwen Image 3.0 by spending its long prompt budget on every panel.
How promptExtend works on Qwen Image 3.0: what the LLM rewrite adds to a short prompt, how the direct and agent modes differ, and why extension and seeds don't mix.
How to prompt Qwen Image 3.0, from sizing inside its pixel budget to layering a scene description, steering with negativePrompt, and locking a result with a seed.
How to use reference images on Qwen Image 3.0 to recolour a packshot, carry a product into a new scene, or stage two products together in one shot.
How to render exact copy with Qwen Image 3.0: quoting the strings you need, building a type hierarchy, setting Chinese and bilingual layouts, and sizing for print.
How to turn structured content into video with Wan 3.0: recipes as cooking videos and public webpages as short landmark or historical montages.
How to use images with Wan 3.0's video generation: pin an image as the first frame, morph between two frames, and carry subject or product identity across new scenes with reference images.
How to write prompts for Wan 3.0 that get you the shot: how the model reads layered directives, camera language, style and register range, directing native audio, and the dimensions and duration rules.
How to use video and audio references with Wan 3.0: carrying a subject across new scenes with referenceVideos, driving video from an audio source with referenceAudios, and combining multiple reference types in one call.
How to compose FLUX 3's synchronized audio: layering ambient and effects, directing music by instrumentation and tempo, directing speech delivery, and rendering multilingual dialogue with clean lip-sync.
How to pin images to specific frame positions in FLUX 3 videos: opening on a source image, storyboarding across positions, ending on a packshot, morphing between two states, and using timestamps for beat-precise timing.
How to build multi-shot video sequences inside one FLUX 3 generation: HARD CUT syntax, shot contrast, time compression, threading an audio bed across the cuts, and pacing shot rhythm.
How to prompt FLUX 3 for text-to-video with synchronized audio: request shape, prompt rewriting, camera language, dimensions, style diversity, and the draft-mode iteration workflow.
How to continue a source clip from its final frames with FLUX 3's video input: single-clip continuation, chaining across multiple generations, deliberate source design, and recovering shots that didn't land.
How to remove objects from images with FLUX Erase, Black Forest Labs's prompt-less mask-driven removal model. What you paint is what disappears.
How to extend an image past its original frame with FLUX Outpainting. No prompt, no mask, just more canvas around what is already there.
How to use FLUX Video Upscale's settings.creativity: precise, source-faithful sharpening for faces and brands, or creative detail restoration for scenery, textures, and crowds.
How to steer FLUX Video Upscale with an optional positivePrompt: name the materials and textures you want rebuilt so the enhancement favours the detail that matters.
How to upscale video with FLUX Video Upscale: the async request, choosing an upscale factor, the input size and length limits, and keeping motion and the original audio.
How to dress a person in any garment from a reference image with FLUX VTO. The call takes one person photo, one garment photo, and a short prompt.
How to edit an existing clip in place with Seedance 2.5 using inputs.video: remove and replace elements, restyle and relight the frame, swap backgrounds, and edit by timestamp.
How to turn a still into video with Seedance 2.5: animate a first frame, interpolate a first and last frame, loop a still, and sequence images as keyframes.
How to structure a 30-second Seedance 2.5 brief as one continuous take or a multi-shot passage, hold continuity across the runtime, and extend past 30 seconds with inputs.video.
How to drive Seedance 2.5 with a reference video for motion, camera, and lip-sync while a reference image supplies identity: motion transfer, character swaps, and clay to finished shot.
How to compose up to 30 images, 10 videos, and 10 audio clips into one Seedance 2.5 call with typed @-tag addressing, from cast ensembles and product families to audio-driven scenes.
How to generate video in 10+ languages with Seedance 2.5: native lip-sync from a text prompt, switching languages within one clip, and controlling accent and delivery.
How to prompt Seedance 2.5 for directed text-to-video: the five-layer shot scaffold, camera vocabulary the model reads directly, second-level timing, and native audio.
How to write prompts for Seedream 4.5, from short interpretive prompts to layered scene descriptions, with legible text rendering at 2K to 4K.
How to prompt Seedream 5.0 Pro for accurate in-image text across 15 native languages, from French posters with accented characters to Arabic, Thai, Korean, and Japanese scripts.
How to prompt Seedream 5.0 Pro for text-to-image, reference-guided editing, and multi-image fusion in commercial workflows: campaign posters, product variants, and editorial still-life.
How to clean up and upscale soft, compressed, AI-generated, or archival footage to 4K with BytePlus Video Enhancement Pro, and when to reach for it.
How to train a custom Exactly Illustrative style model on a brand's visual identity, then use it for consistent text-to-image and image-to-image generation in that look.
How to control vocal delivery in Fish Audio S2-Pro with bracket tags. The tag system steers emotion, expression, paralanguage, and phoneme-level pronunciation in one inline syntax.
How to generate two-speaker dialogue audio in a single request to Fish Audio S2-Pro using inline speaker tags. One call, two voices, full per-speaker emotion control.
How to transform game assets while preserving their structure using Canny edge detection, ControlNet, and LoRAs for consistent style across variations.
How to edit an existing clip with Gemini Omni Flash 1.1: relighting, weather, and restyling through inputs.video, why short prompts win, and how to chain edits.
How to extend a clip past the 10-second ceiling with Gemini Omni Flash 1.1: sending duration with inputs.video, prompting continuation, and chaining extensions.
How to pin the opening and closing frames of a Gemini Omni Flash 1.1 clip with inputs.frameImages, and prompt the transition the model builds between them.
How to prompt Gemini Omni Flash 1.1 for video: the five-element structure, camera vocabulary, holding a single shot, directing audio, and the sampling controls.
How to carry a character, product, or style into a Gemini Omni Flash 1.1 shot with inputs.referenceImages and inputs.referenceVideos, and when to combine the two.
How to size and time a Gemini Omni Flash 1.1 video: the four resolution presets, the supported width and height pairs, and how tier and duration drive what a clip costs.
How to prompt Gemini Omni Flash for cinematic video using Google's five-element structure, camera language, and the less-prescriptive sweet spot.
How to edit existing footage with Gemini Omni Flash's inputs.video parameter to relight, restyle, swap weather, or add characters while preserving the source's composition and motion.
How to use Gemini Omni Flash's reference image workflow to lock a visual style, hold a character across scenes, or guide a video through storyboard key beats.
How to use Nano Banana 2 reference images to keep the same character or product identical across new scenes and styles.
How to generate images from real-time information with Nano Banana 2 using web and image search grounding via providerSettings.google.
How to merge several reference images, a product, a subject, a backdrop, or a style, into a single coherent image with Nano Banana 2.
How to write prompts for Nano Banana 2: detailed scene descriptions, structured layering, legible text rendering, and thinking-level control.
How to compose what surrounds the avatar in Avatar V output: the background, fit mode, aspect ratio for the target platform, and burned-in captions.
How to choose between Avatar V's two input modes: generate the voice from a script, or drive the avatar with your own recorded audio.
How to generate 3D models with Rodin Gen-2 from a text prompt or reference images: choosing inputs, using multi-view, prompting for 3D form, and working with the GLB output.
How to match Rodin Gen-2 output to your pipeline: Quad vs Raw topology, quality presets vs custom polygon counts, HighPack 4K textures, T/A pose for rigging, and PBR vs baked materials.
How Ideogram 4.0's two prompting modes work: natural language with Magic Prompt expansion for quick exploration, and the full trained JSON schema for explicit per-element control.
How to use Ideogram 4.0 for typography-heavy design where the text has to be readable and exactly right, the layout has to land, and the palette has to lock to brand.
How to erase objects, people, watermarks, and timestamps from an image with Ideogram Object Remover using a mask, with no prompt required.
How to write LLM system prompts that produce text TTS-2 can synthesize naturally, with normalization, filler words, and emphasis cues handled before the audio call.
How to use natural-language steering tags to control emotion, pacing, volume, and vocal style in TTS-2 speech output.
How to use Kling 3.0 Turbo's inline shot-list syntax to direct multi-shot video reels in a single API call, with shot-by-shot timing and prompt control.
How to steer Krea 2 with the creativity parameter, weighted style reference images, and moodboards. Together these cover faithful renders through bold reinterpretation.
How to drive a lip-synced performance with LTX-2.5 Pro: pairing inputs.audio with a reference image, matching voice to subject, and the length, resolution, and framing rules.
How to control the camera in LTX-2.5 Pro: the eight settings.cameraMovement presets, the prompt camera vocabulary, and when to reach for a reliable preset versus a described move.
How to animate images with LTX-2.5 Pro using inputs.frameImages: turning a still into motion, directing the ending with a last frame, and building clean loops.
How to prompt LTX-2.5 Pro for multi-shot video: cutting between shots in one generation, naming transitions, holding character and audio across cuts, and giving each shot a job.
How to generate synchronized audio with LTX-2.5 Pro: turning on settings.audio, prompting ambient sound and effects, directing spoken dialogue with accent and lip-sync, and balancing the mix.
How to write text-to-video prompts for LTX-2.5 Pro: the six-part shot scaffold, directing the action and camera, matching detail to shot scale, and prompting its native audio.
How to transform footage with Luma Ray 3.2: restyle or reskin a clip while its motion carries through, using the strength dial and per-signal conditioning controls.
How to generate cinematic video with Luma Ray 3.2: text-to-video, image-to-video, frame-level keyframes, and the resolution, duration, HDR, and loop controls.
How to reframe footage with Luma Ray 3.2: convert a clip to a new aspect ratio or a larger canvas while the model extends the scene, using inputs.video and sourcePosition.
How to generate 3D models with Meshy-6 from a prompt or photos: choosing text vs image input, framing a clean reference, using multiple views, and prompting for 3D form.
How to edit an image with Muse Image: passing one source image, targeting a region in words, removing objects, restoring old photos, and refining a result across passes.
How to ground Muse Image in real facts and real references with settings.webSearch and settings.imageSearch, and when to switch the lookups off.
How to build data-accurate charts and infographics with Muse Image: what settings.shell does, how to write your figures into the prompt, and when to switch it off.
How to build one image out of up to ten reference images with Muse Image: giving each reference a job in the prompt, and sizing the output with resolution 2K.
How to prompt Muse Image: writing a brief it can plan a layout from, setting the thinking level, picking one of the eight size pairs, and working without a seed.
How to render exact, legible copy inside an image with Muse Image: quoting the literal string, keeping it short, directing the type, and landing multi-element layouts.
How to generate video with MiniMax H3 Max: the fixed size pairs at 480p and 768p, animating a first and last frame, native synced audio, and rendering faster than playback.
How to edit a finished clip with a prompt in MiniMax H3: replace a subject, relight a scene, add or remove elements, and combine edits, keeping the rest untouched.
How to animate a still image with MiniMax H3 and bridge a first and last frame into one continuous shot, with the output following the image's own aspect ratio.
How to generate video from text with MiniMax H3: the six 2K aspect ratios, 5 to 15 second durations, native synced audio, and prompting for cinematic shots.
How to transfer motion, a camera move, or an acting performance from a reference video onto a new subject with MiniMax H3 Omni Reference.
How to lock a character, product, or style across a new MiniMax H3 shot with Omni Reference images, and address each reference by index in the prompt.
How to direct MiniMax H3's native audio from the prompt: ambience, synced sound effects, music, and spoken dialogue with lip-sync.
How to give Kimi K2.6 access to your own functions through Runware's OpenAI-compatible endpoint: defining tools, running the call loop, making parallel tool calls, and controlling tool selection.
How to write prompts for GPT Image 2 across the cases the model handles unusually well: photorealism, accurate text, world-knowledge composition, and multi-image editing workflows.
How to use PixVerse V6 to generate multi-shot video reels from one or two anchor frames in a single call, using the beat-structured prompt pattern.
How to prompt P-Image-Ideogram in both modes: natural language for quick exploration and structured JSON for explicit control over layout, colour, and typography.
How to design posters, packaging, ads, and menus with P-Image-Ideogram, and the workflow that gets on-image copy to render clean and exact.
How to dress a person in a full outfit with Pruna P-Image-Try-On, passing each garment as its own reference and matching a target pose.
How to use Pruna P-Video-Animate to bring a still reference image to life by inheriting the motion, timing, and camera move from a source video.
How to swap a single on-camera object (a product or a garment) in a source video with Pruna P-Video-Replace, without touching the rest of the frame.
How to recreate iconic film scenes with Bytedance Seedance, then recast the on-camera character with Pruna P-Video-Replace to drop yourself or any reference into the shot.
How to use Pruna P-Video-Replace to swap the on-camera character in an existing video with one from a reference image while preserving the original motion, timing, camera, lighting, and audio.
How to use the settings.colors and settings.backgroundColor parameters to lock generated images to a specific color palette or background color.
How to write effective prompts for Recraft V4.1, from short interpretive prompts to structured multi-layer descriptions for precise creative control.
How to pick between Recraft V4 Styles match modes: precise for strict style adherence, flexible for a looser vibe, with side-by-side comparisons.
How to prompt Recraft V4 Styles: how references pin the style, what the prompt controls, dimensions, and patterns for consistent output.
How to build a reusable style in Recraft V4 Styles: reference set strategy, styleId reuse across many generations, and same-family portability across the raster and vector tiers.
How to make localised edits to existing footage with Runway Aleph 2 that change only the targeted region and leave the rest of the clip untouched.
How to write Runway Gen-4.5 image-to-video prompts that direct motion instead of redescribing the scene, using the camera and subject channels.
How to generate consistent and professional sticker collections using specialized diffusion models and LoRAs.
How to use SDXL's refiner pipeline to enhance fine details and textures in the final denoising pass.
How to use scoringPrompt and scoringRubric on Sourceful Riverflow 2.5 Pro to drive different production workflows from the same brand inputs.
How to control how much detail Topaz Wonder 3.5 rebuilds with settings.enhancementStrength: what low, medium, and high change, and when to dial it down for faces or up for texture.
How to add film grain to a Wonder 3.5 upscale with settings.grain: the silver, gaussian, and grey grain models and the strength, density, and size controls.
How to restore old and degraded photos with Topaz Wonder 3.5: faded scans, over-compressed uploads, noisy low-light shots, and tiny legacy files rebuilt as they upscale.
How to sharpen text, charts, tables, and packaging with Topaz Wonder 3.5, recovering legible type and clean structured graphics from soft or compressed sources.
How to upscale images with Topaz Wonder 3.5: the request shape, choosing an upscale factor, the input and output size limits, and what generative upscaling recovers.
How to edit an image with Grok Imagine Image 2.0: recolour, restyle, remove objects, swap backgrounds, and relight from one reference image and a prompt.
How to write text-to-image prompts for Grok Imagine Image 2.0: structuring a shot, directing the camera and light, the 1K and 2K aspect pairs, and its factual detail.
How to render accurate, readable text with Grok Imagine Image 2.0: quoting the exact words, directing typography and hierarchy, and holding up on small, dense layouts.
How to generate images with accurate, readable text using xAI Grok Imagine. Prompt for the text content, the placement, and the script you want rendered.