T2V · I2V / FL2VA · R2V / Ref2VA · 4–15 seconds
A strong MiniMax H3 workflow begins with the type of control your shot needs. Start from text for an open canvas, use first or last frames for composition, or assign image, video, and audio references for identity, movement, style, and voice.
Independent service. Not affiliated with, endorsed by, or sponsored by MiniMax or any model owner.
MiniMax H3 workflow / production guide August 7, 2026
The MiniMax H3 workflow has three practical starting paths: T2V for text-only direction, I2V/FL2VA for first and last frames, and R2V/Ref2VA for multimodal references. The best path is the one that expresses the shot’s hardest constraint with the least contradictory context. Choosing the wrong mode creates prompt clutter, unnecessary uploads, and false expectations that cannot be repaired by adding more adjectives.
Text-to-video uses the FL2VA base with no image connected. It is the cleanest workflow for ideation, prompt benchmarking, atmospheric inserts, and shots where subject placement can be invented. Write the overall scene first, then timed beats: who or what is present, what changes, camera position and movement, lighting, physical reactions, dialogue, effects, and music. Put requirements in production order rather than grouping random style words.

T2V also exposes whether the model understood your language because no input frame can rescue the composition. Keep the first test short and economical, save the exact prompt and settings, and compare prompt adherence across several candidates. A beautiful thumbnail with unusable motion is not the winner; judge anatomy, object permanence, camera logic, timing, and synchronized sound through the whole clip.
Image-to-video uses the FL2VA path with a first frame, a last frame, or both. One first frame fixes the opening composition; one last frame defines a destination; two frames ask the model to create a plausible transition. Describe change rather than redundantly cataloging the still image: camera travel, body movement, material response, environmental motion, timing, sound, and the details that must remain stable.

A keyframe is strong control over composition, not a guarantee that every identity detail survives fifteen seconds. Use clean, correctly oriented images with the intended crop and enough visual information for motion. If the real need is to borrow identity from one image, motion from a video, and voice from audio, stop forcing those roles into a two-frame workflow and move to Ref2VA.
Reference-to-video uses separate Ref2VA weights. It can combine up to nine images, three videos, and three audio clips, with twelve files total. Video and audio each have published per-file and total-duration limits. Refer to inputs in connection order and assign every source a role: `<Picture 1>` defines face and wardrobe, `<Video 1>` supplies camera rhythm, and `<Audio 1>` provides ambience or voice texture.

Explicit roles matter because two references may disagree. A fashion still might demand a fixed silhouette while a motion clip twists the body; an audio beat may require a cut that conflicts with a last-frame destination. Remove redundant files, decide which property wins, and write retention priorities. The model’s multimodal capacity is not permission to upload everything you have.
Choose 4–15 seconds, an aspect ratio appropriate to delivery, and 768P for exploration or 2K when fine detail justifies the cost. Native stereo audio is generated with the image sequence, so the prompt should direct room tone, physical effects, dialogue timing, and music instead of assuming sound will be repaired later. Always review pronunciation, lip sync, clipping, unwanted music, and channel balance.

Log the input path, prompt, files, duration, resolution, aspect ratio, seed or run identifier, credits, and a short rejection reason. Change one high-value variable for the next take. This turns H3 from a slot machine into an evaluable production process. When a shot fails for composition, do not jump to higher resolution; when it fails for reference conflict, do not add another reference.
Nothing is visually fixed. Choose T2V and test whether the written direction alone produces a usable composition, performance, camera path, and soundscape.
The first or last composition is fixed. Choose I2V/FL2VA, upload one or two keyframes, and prompt the temporal change between those anchors.
Identity, motion, style, or voice comes from files. Choose R2V/Ref2VA, label every source, assign one job to each, and remove conflicts before generation.
Text opens composition. Keyframes anchor composition. Multimodal references assign identity, motion, style, camera, or audio to concrete sources.
Select T2V, I2V/FL2VA, or R2V/Ref2VA based on the hardest constraint. Do not mix keyframe and reference strategies without a supported contract.
Write timed action, camera, continuity, and sound. Label reference inputs in order and state exactly which property each source controls.
Review the full clip, log failures, keep stable settings, and change one variable. Optimize for edit-ready seconds rather than the most attractive still frame.
Start with the smallest input set that expresses the brief, select 4–15 seconds and 768P or 2K, then review the complete result.
Credits vary with output resolution, duration, reference-video seconds, and additional images. Prove the direction before scaling settings.
Billed $238.80 yearly
Billed $418.80 yearly
Billed $898.80 yearly
Included models
Credits activate instantly for MiniMax H3 generation workflows.

Give every prompt sentence and reference file one purpose. Remove anything that cannot win a conflict or change the target shot.
Build the First TakeChoose T2V when composition is open, I2V/FL2VA when a first frame, last frame, or both must anchor the shot, and R2V/Ref2VA when identity, style, motion, camera, or voice comes from multiple reference files. Use the least complex path that expresses the hard constraint.
FL2VA uses text plus zero, one, or two keyframes for text-to-video and first/last-frame generation. Ref2VA uses different diffusion weights for multimodal reference generation from images, videos, and audio. A downloaded FL2VA model does not automatically satisfy an R2V workflow.
A first frame fixes the opening composition, a last frame defines the destination, and two frames guide the transition between them. Prompt the movement, camera, timing, environmental response, sound, and invariants. Do not waste the prompt merely restating every visible detail in the images.
Reference inputs should be addressed in connection order, such as `<Picture 1>`, `<Video 1>`, and `<Audio 1>`. State one explicit role for each source—identity, wardrobe, style, camera movement, performance timing, ambience, or voice—and resolve conflicts before generation.
The released H3 specification supports 4–15 second output at 24 FPS. Local ComfyUI duration controls may snap to the model’s compatible frame grid. Longer or larger clips increase compute and memory, so validate the shot with a shorter preview before scaling.
Yes. H3 jointly generates video and 32 kHz stereo audio rather than adding a stock track afterward. Direct dialogue, effects, music, distance, and timing in the prompt, then review pronunciation, synchronization, clipping, unwanted sound, and usage rights before publishing.