
Text to image
The foundational pipeline: generate images from text prompts using diffusion models, with control over every step of the process.
Understand core generative AI concepts and explore model-specific guides with practical examples and code.
Deep dives into every parameter and feature of the image generation pipeline. Each page covers how a feature works, its visual impact, and practical usage patterns with interactive examples.
Step-by-step tutorials built around specific models. Each guide walks through a real use case with code, prompts, and results you can reproduce.
How to compose FLUX 3's synchronized audio: layering ambient and effects, directing music by instrumentation and tempo, directing speech delivery, and rendering multilingual dialogue with clean lip-sync.

How to build multi-shot video sequences inside one FLUX 3 generation: HARD CUT syntax, shot contrast, time compression, threading an audio bed across the cuts, and pacing shot rhythm.

How to edit a finished clip with a prompt in MiniMax H3: replace a subject, relight a scene, add or remove elements, and combine edits, keeping the rest untouched.

How to transfer motion, a camera move, or an acting performance from a reference video onto a new subject with MiniMax H3 Omni Reference.

How to use reference images on Qwen Image 3.0 to recolour a packshot, carry a product into a new scene, or stage two products together in one shot.

How to design posters, packaging, ads, and menus with P-Image-Ideogram, and the workflow that gets on-image copy to render clean and exact.

How to write multi-shot storyboard prompts for HappyHorse 1.1 to direct cinematic sequences with shot sizes, camera movement, and subject continuity in a single call.
How to use HappyHorse 1.1's reference workflow to cast one or more characters into a generated video and preserve their identity through every cut.
How to generate infographics, comparison grids, newspaper pages, and interface mockups with Qwen Image 3.0 by spending its long prompt budget on every panel.
How promptExtend works on Qwen Image 3.0: what the LLM rewrite adds to a short prompt, how the direct and agent modes differ, and why extension and seeds don't mix.
How to prompt Qwen Image 3.0, from sizing inside its pixel budget to layering a scene description, steering with negativePrompt, and locking a result with a seed.
How to use reference images on Qwen Image 3.0 to recolour a packshot, carry a product into a new scene, or stage two products together in one shot.
How to render exact copy with Qwen Image 3.0: quoting the strings you need, building a type hierarchy, setting Chinese and bilingual layouts, and sizing for print.
How to compose FLUX 3's synchronized audio: layering ambient and effects, directing music by instrumentation and tempo, directing speech delivery, and rendering multilingual dialogue with clean lip-sync.
How to pin images to specific frame positions in FLUX 3 videos: opening on a source image, storyboarding across positions, ending on a packshot, morphing between two states, and using timestamps for beat-precise timing.
How to build multi-shot video sequences inside one FLUX 3 generation: HARD CUT syntax, shot contrast, time compression, threading an audio bed across the cuts, and pacing shot rhythm.
How to prompt FLUX 3 for text-to-video with synchronized audio: request shape, prompt rewriting, camera language, dimensions, style diversity, and the draft-mode iteration workflow.
How to continue a source clip from its final frames with FLUX 3's video input: single-clip continuation, chaining across multiple generations, deliberate source design, and recovering shots that didn't land.
How to remove objects from images with FLUX Erase, Black Forest Labs's prompt-less mask-driven removal model. What you paint is what disappears.
How to extend an image past its original frame with FLUX Outpainting. No prompt, no mask, just more canvas around what is already there.
How to dress a person in any garment from a reference image with FLUX VTO. The call takes one person photo, one garment photo, and a short prompt.
How to edit an existing clip in place with Seedance 2.5 using inputs.video: remove and replace elements, restyle and relight the frame, swap backgrounds, and edit by timestamp.
How to turn a still into video with Seedance 2.5: animate a first frame, interpolate a first and last frame, loop a still, and sequence images as keyframes.
How to structure a 30-second Seedance 2.5 brief as one continuous take or a multi-shot passage, hold continuity across the runtime, and extend past 30 seconds with inputs.video.
How to drive Seedance 2.5 with a reference video for motion, camera, and lip-sync while a reference image supplies identity: motion transfer, character swaps, and clay to finished shot.
How to compose up to 30 images, 10 videos, and 10 audio clips into one Seedance 2.5 call with typed @-tag addressing, from cast ensembles and product families to audio-driven scenes.
How to generate video in 10+ languages with Seedance 2.5: native lip-sync from a text prompt, switching languages within one clip, and controlling accent and delivery.
How to prompt Seedance 2.5 for directed text-to-video: the five-layer shot scaffold, camera vocabulary the model reads directly, second-level timing, and native audio.
How to write prompts for Seedream 4.5, from short interpretive prompts to layered scene descriptions, with legible text rendering at 2K to 4K.
How to prompt Seedream 5.0 Pro for accurate in-image text across 15 native languages, from French posters with accented characters to Arabic, Thai, Korean, and Japanese scripts.
How to prompt Seedream 5.0 Pro for text-to-image, reference-guided editing, and multi-image fusion in commercial workflows: campaign posters, product variants, and editorial still-life.
How to clean up and upscale soft, compressed, AI-generated, or archival footage to 4K with BytePlus Video Enhancement Pro, and when to reach for it.
How to train a custom Exactly Illustrative style model on a brand's visual identity, then use it for consistent text-to-image and image-to-image generation in that look.
How to control vocal delivery in Fish Audio S2-Pro with bracket tags. The tag system steers emotion, expression, paralanguage, and phoneme-level pronunciation in one inline syntax.
How to generate two-speaker dialogue audio in a single request to Fish Audio S2-Pro using inline speaker tags. One call, two voices, full per-speaker emotion control.
How to transform game assets while preserving their structure using Canny edge detection, ControlNet, and LoRAs for consistent style across variations.
How to prompt Gemini Omni Flash for cinematic video using Google's five-element structure, camera language, and the less-prescriptive sweet spot.
How to edit existing footage with Gemini Omni Flash's inputs.video parameter to relight, restyle, swap weather, or add characters while preserving the source's composition and motion.
How to use Gemini Omni Flash's reference image workflow to lock a visual style, hold a character across scenes, or guide a video through storyboard key beats.
How to use Nano Banana 2 reference images to keep the same character or product identical across new scenes and styles.
How to generate images from real-time information with Nano Banana 2 using web and image search grounding via providerSettings.google.
How to merge several reference images, a product, a subject, a backdrop, or a style, into a single coherent image with Nano Banana 2.
How to write prompts for Nano Banana 2: detailed scene descriptions, structured layering, legible text rendering, and thinking-level control.
How to compose what surrounds the avatar in Avatar V output: the background, fit mode, aspect ratio for the target platform, and burned-in captions.
How to choose between Avatar V's two input modes: generate the voice from a script, or drive the avatar with your own recorded audio.
How to generate 3D models with Rodin Gen-2 from a text prompt or reference images: choosing inputs, using multi-view, prompting for 3D form, and working with the GLB output.
How to match Rodin Gen-2 output to your pipeline: Quad vs Raw topology, quality presets vs custom polygon counts, HighPack 4K textures, T/A pose for rigging, and PBR vs baked materials.
How Ideogram 4.0's two prompting modes work: natural language with Magic Prompt expansion for quick exploration, and the full trained JSON schema for explicit per-element control.
How to use Ideogram 4.0 for typography-heavy design where the text has to be readable and exactly right, the layout has to land, and the palette has to lock to brand.
How to erase objects, people, watermarks, and timestamps from an image with Ideogram Object Remover using a mask, with no prompt required.
How to write LLM system prompts that produce text TTS-2 can synthesize naturally, with normalization, filler words, and emphasis cues handled before the audio call.
How to use natural-language steering tags to control emotion, pacing, volume, and vocal style in TTS-2 speech output.
How to use Kling 3.0 Turbo's inline shot-list syntax to direct multi-shot video reels in a single API call, with shot-by-shot timing and prompt control.
How to steer Krea 2 with the creativity parameter, weighted style reference images, and moodboards. Together these cover faithful renders through bold reinterpretation.
How to drive a lip-synced performance with LTX-2.5 Pro: pairing inputs.audio with a reference image, matching voice to subject, and the length, resolution, and framing rules.
How to control the camera in LTX-2.5 Pro: the eight settings.cameraMovement presets, the prompt camera vocabulary, and when to reach for a reliable preset versus a described move.
How to animate images with LTX-2.5 Pro using inputs.frameImages: turning a still into motion, directing the ending with a last frame, and building clean loops.
How to prompt LTX-2.5 Pro for multi-shot video: cutting between shots in one generation, naming transitions, holding character and audio across cuts, and giving each shot a job.
How to generate synchronized audio with LTX-2.5 Pro: turning on settings.audio, prompting ambient sound and effects, directing spoken dialogue with accent and lip-sync, and balancing the mix.
How to write text-to-video prompts for LTX-2.5 Pro: the six-part shot scaffold, directing the action and camera, matching detail to shot scale, and prompting its native audio.
How to transform footage with Luma Ray 3.2: restyle or reskin a clip while its motion carries through, using the strength dial and per-signal conditioning controls.
How to generate cinematic video with Luma Ray 3.2: text-to-video, image-to-video, frame-level keyframes, and the resolution, duration, HDR, and loop controls.
How to reframe footage with Luma Ray 3.2: convert a clip to a new aspect ratio or a larger canvas while the model extends the scene, using inputs.video and sourcePosition.
How to generate 3D models with Meshy-6 from a prompt or photos: choosing text vs image input, framing a clean reference, using multiple views, and prompting for 3D form.
How to edit a finished clip with a prompt in MiniMax H3: replace a subject, relight a scene, add or remove elements, and combine edits, keeping the rest untouched.
How to animate a still image with MiniMax H3 and bridge a first and last frame into one continuous shot, with the output following the image's own aspect ratio.
How to generate video from text with MiniMax H3: the six 2K aspect ratios, 5 to 15 second durations, native synced audio, and prompting for cinematic shots.
How to transfer motion, a camera move, or an acting performance from a reference video onto a new subject with MiniMax H3 Omni Reference.
How to lock a character, product, or style across a new MiniMax H3 shot with Omni Reference images, and address each reference by index in the prompt.
How to direct MiniMax H3's native audio from the prompt: ambience, synced sound effects, music, and spoken dialogue with lip-sync.
How to give Kimi K2.6 access to your own functions through Runware's OpenAI-compatible endpoint: defining tools, running the call loop, making parallel tool calls, and controlling tool selection.
How to write prompts for GPT Image 2 across the cases the model handles unusually well: photorealism, accurate text, world-knowledge composition, and multi-image editing workflows.
How to use PixVerse V6 to generate multi-shot video reels from one or two anchor frames in a single call, using the beat-structured prompt pattern.
How to prompt P-Image-Ideogram in both modes: natural language for quick exploration and structured JSON for explicit control over layout, colour, and typography.
How to design posters, packaging, ads, and menus with P-Image-Ideogram, and the workflow that gets on-image copy to render clean and exact.
How to dress a person in a full outfit with Pruna P-Image-Try-On, passing each garment as its own reference and matching a target pose.
How to use Pruna P-Video-Animate to bring a still reference image to life by inheriting the motion, timing, and camera move from a source video.
How to swap a single on-camera object (a product or a garment) in a source video with Pruna P-Video-Replace, without touching the rest of the frame.
How to recreate iconic film scenes with Bytedance Seedance, then recast the on-camera character with Pruna P-Video-Replace to drop yourself or any reference into the shot.
How to use Pruna P-Video-Replace to swap the on-camera character in an existing video with one from a reference image while preserving the original motion, timing, camera, lighting, and audio.
How to use the settings.colors and settings.backgroundColor parameters to lock generated images to a specific color palette or background color.
How to write effective prompts for Recraft V4.1, from short interpretive prompts to structured multi-layer descriptions for precise creative control.
How to place text precisely inside a Reve 2.1 composition and render accurate copy across Chinese, Devanagari, Cyrillic, Arabic, and mixed-script layouts at 4K.
How to use settings.postprocessing on Reve 2.1 to run a pipeline over the model output, and when to reach for removeBackground for cutout deliverables.
How to prompt Reve 2.1 for native 4K commercial imagery, from picking among seventeen preset aspect ratios to directing dense layout-aware compositions.
How to use Reve 2.1 referenceImages with in-prompt frame tags to edit an existing asset or remix a subject into a new scene, from wall colour swaps to multi-product composites.
How to make localised edits to existing footage with Runway Aleph 2 that change only the targeted region and leave the rest of the clip untouched.
How to write Runway Gen-4.5 image-to-video prompts that direct motion instead of redescribing the scene, using the camera and subject channels.
How to generate consistent and professional sticker collections using specialized diffusion models and LoRAs.
How to use SDXL's refiner pipeline to enhance fine details and textures in the final denoising pass.
How to use scoringPrompt and scoringRubric on Sourceful Riverflow 2.5 Pro to drive different production workflows from the same brand inputs.
How to control how much detail Topaz Wonder 3.5 rebuilds with settings.enhancementStrength: what low, medium, and high change, and when to dial it down for faces or up for texture.
How to add film grain to a Wonder 3.5 upscale with settings.grain: the silver, gaussian, and grey grain models and the strength, density, and size controls.
How to restore old and degraded photos with Topaz Wonder 3.5: faded scans, over-compressed uploads, noisy low-light shots, and tiny legacy files rebuilt as they upscale.
How to sharpen text, charts, tables, and packaging with Topaz Wonder 3.5, recovering legible type and clean structured graphics from soft or compressed sources.
How to upscale images with Topaz Wonder 3.5: the request shape, choosing an upscale factor, the input and output size limits, and what generative upscaling recovers.
How to edit an image with Grok Imagine Image 2.0: recolour, restyle, remove objects, swap backgrounds, and relight from one reference image and a prompt.
How to write text-to-image prompts for Grok Imagine Image 2.0: structuring a shot, directing the camera and light, the 1K and 2K aspect pairs, and its factual detail.
How to render accurate, readable text with Grok Imagine Image 2.0: quoting the exact words, directing typography and hierarchy, and holding up on small, dense layouts.
How to generate images with accurate, readable text using xAI Grok Imagine. Prompt for the text content, the placement, and the script you want rendered.