Home
Create
Image
Video
Creations
Saved
Upgrade

MiniMax H3 Workflow Guide

A strong MiniMax H3 workflow begins with the type of control your shot needs. Start from text for an open canvas, use first or last frames for composition, or assign image, video, and audio references for identity, movement, style, and voice.

142 Credits

MiniMax H3 workflow / production guide August 7, 2026

Choose the input contract before writing the prompt

The MiniMax H3 workflow has three practical starting paths: T2V for text-only direction, I2V/FL2VA for first and last frames, and R2V/Ref2VA for multimodal references. The best path is the one that expresses the shot’s hardest constraint with the least contradictory context. Choosing the wrong mode creates prompt clutter, unnecessary uploads, and false expectations that cannot be repaired by adding more adjectives.

Use T2V when the composition is still open

Text-to-video uses the FL2VA base with no image connected. It is the cleanest workflow for ideation, prompt benchmarking, atmospheric inserts, and shots where subject placement can be invented. Write the overall scene first, then timed beats: who or what is present, what changes, camera position and movement, lighting, physical reactions, dialogue, effects, and music. Put requirements in production order rather than grouping random style words.

T2V also exposes whether the model understood your language because no input frame can rescue the composition. Keep the first test short and economical, save the exact prompt and settings, and compare prompt adherence across several candidates. A beautiful thumbnail with unusable motion is not the winner; judge anatomy, object permanence, camera logic, timing, and synchronized sound through the whole clip.

Use I2V or FL2VA to anchor start and finish

Image-to-video uses the FL2VA path with a first frame, a last frame, or both. One first frame fixes the opening composition; one last frame defines a destination; two frames ask the model to create a plausible transition. Describe change rather than redundantly cataloging the still image: camera travel, body movement, material response, environmental motion, timing, sound, and the details that must remain stable.

A keyframe is strong control over composition, not a guarantee that every identity detail survives fifteen seconds. Use clean, correctly oriented images with the intended crop and enough visual information for motion. If the real need is to borrow identity from one image, motion from a video, and voice from audio, stop forcing those roles into a two-frame workflow and move to Ref2VA.

Use R2V or Ref2VA when references have different jobs

Reference-to-video uses separate Ref2VA weights. It can combine up to nine images, three videos, and three audio clips, with twelve files total. Video and audio each have published per-file and total-duration limits. Refer to inputs in connection order and assign every source a role: `<Picture 1>` defines face and wardrobe, `<Video 1>` supplies camera rhythm, and `<Audio 1>` provides ambience or voice texture.

Explicit roles matter because two references may disagree. A fashion still might demand a fixed silhouette while a motion clip twists the body; an audio beat may require a cut that conflicts with a last-frame destination. Remove redundant files, decide which property wins, and write retention priorities. The model’s multimodal capacity is not permission to upload everything you have.

Build a review loop around usable seconds

Choose 4–15 seconds, an aspect ratio appropriate to delivery, and 768P for exploration or 2K when fine detail justifies the cost. Native stereo audio is generated with the image sequence, so the prompt should direct room tone, physical effects, dialogue timing, and music instead of assuming sound will be repaired later. Always review pronunciation, lip sync, clipping, unwanted music, and channel balance.

Log the input path, prompt, files, duration, resolution, aspect ratio, seed or run identifier, credits, and a short rejection reason. Change one high-value variable for the next take. This turns H3 from a slot machine into an evaluable production process. When a shot fails for composition, do not jump to higher resolution; when it fails for reference conflict, do not add another reference.

A simple H3 workflow decision tree

Nothing is visually fixed. Choose T2V and test whether the written direction alone produces a usable composition, performance, camera path, and soundscape.

The first or last composition is fixed. Choose I2V/FL2VA, upload one or two keyframes, and prompt the temporal change between those anchors.

Identity, motion, style, or voice comes from files. Choose R2V/Ref2VA, label every source, assign one job to each, and remove conflicts before generation.

Explore the implementation details

01 — Workflow modes

Give the hardest constraint the right input type.

Text opens composition. Keyframes anchor composition. Multimodal references assign identity, motion, style, camera, or audio to concrete sources.

03 — Production loop

Choose, direct, then evaluate.

01 / 03
Choose

Select T2V, I2V/FL2VA, or R2V/Ref2VA based on the hardest constraint. Do not mix keyframe and reference strategies without a supported contract.

02 / 03
Direct

Write timed action, camera, continuity, and sound. Label reference inputs in order and state exactly which property each source controls.

03 / 03
Evaluate

Review the full clip, log failures, keep stable settings, and change one variable. Optimize for edit-ready seconds rather than the most attractive still frame.

04 — Create

Choose the workflow above and direct one controlled shot.

Start with the smallest input set that expresses the brief, select 4–15 seconds and 768P or 2K, then review the complete result.

05 — Pricing

Price the workflow you will actually run

Credits vary with output resolution, duration, reference-video seconds, and additional images. Prove the direction before scaling settings.

Cancel anytime

Standard

$39.90$19.90/ mo

Billed $238.80 yearly

Save $240/year - 50% Off
9,360
credits
$0.11/s
Lowest cost per second
  • Commercial Usage RightsYearly Only
  • Up to 624 videos
  • Video models low to $0.11/s
  • Up to 9,360 images
  • All AI Models
  • Priority Support
  • Priority Processing Speed
  • Remove Watermark
  • Free Video UpscalerYearly Only
  • Free Frame InterpolationYearly Only
  • More Free AI ToolsYearly Only
Most Popular

Business

$69.90$34.90/ mo

Billed $418.80 yearly

Save $420/year - 55% Off
20,400
credits
$0.09/s
Lowest cost per second
  • Commercial Usage RightsYearly Only
  • Up to 1,360 videos
  • Video models low to $0.09/s
  • Up to 20,400 images
  • All AI Models
  • Priority Support
  • Dedicated Support
  • Priority Processing Speed
  • Remove Watermark
  • Free Video UpscalerYearly Only
  • Free Frame InterpolationYearly Only
  • More Free AI ToolsYearly Only
Best Value

Enterprise

$124.90$74.90/ mo

Billed $898.80 yearly

Save $600/year - 60% Off
Quantity Adjustment2x
5x Best Value
58,800
credits
$0.06/s
Lowest cost per second
  • Commercial Usage RightsYearly Only
  • Up to 3,920 videos
  • Video models low to $0.06/s
  • Up to 58,800 images
  • All AI Models
  • Priority Support
  • Dedicated Support
  • Priority Processing Speed
  • Remove Watermark
  • Free Video UpscalerYearly Only
  • Free Frame InterpolationYearly Only
  • More Free AI ToolsYearly Only

Included models

Ideogram 4
Z-Image Turbo
Grok Imagine Image
Nano Banana
Nano Banana Pro
Nano Banana 2
GPT Image 2
Seedream 5.0 Lite
Seedream 4.5
MiniMax H3
Seedance 2.0
Seedance 2.0 Fast
Seedance 2.0 Mini
Seedance 2.5Coming SoonGet Early Bird
Seedance 1.5 Pro
Grok Imagine Video
Gemini Omni Flash
Veo 3.1 Quality
Veo 3.1 Lite
Veo 3.1 Fast
Wan 2.7
Wan 2.6
Wan 2.5
Secure checkoutstripe

Credits activate instantly for MiniMax H3 generation workflows.

Free trial credits on signupInstant credit deliveryMiniMax H3 and full model access
Supported payment methods
Mastercard
VISAVisa
AMEXAmerican Express
Apple Pay
Google Pay
linkLink
UnionPay
JCBJCB
DISCOVERDiscover
SEPASEPA
Creator planning a MiniMax H3 T2V I2V or R2V workflow
Workflow rule

The fewest clear inputs usually beat a crowded brief.

Give every prompt sentence and reference file one purpose. Remove anything that cannot win a conflict or change the target shot.

Build the First Take
06 — Workflow questions

MiniMax H3 Workflow FAQ

Which MiniMax H3 workflow should I choose?

Choose T2V when composition is open, I2V/FL2VA when a first frame, last frame, or both must anchor the shot, and R2V/Ref2VA when identity, style, motion, camera, or voice comes from multiple reference files. Use the least complex path that expresses the hard constraint.

What is the difference between FL2VA and Ref2VA?

FL2VA uses text plus zero, one, or two keyframes for text-to-video and first/last-frame generation. Ref2VA uses different diffusion weights for multimodal reference generation from images, videos, and audio. A downloaded FL2VA model does not automatically satisfy an R2V workflow.

How do first and last frames work in MiniMax H3?

A first frame fixes the opening composition, a last frame defines the destination, and two frames guide the transition between them. Prompt the movement, camera, timing, environmental response, sound, and invariants. Do not waste the prompt merely restating every visible detail in the images.

How should I label MiniMax H3 reference inputs?

Reference inputs should be addressed in connection order, such as `<Picture 1>`, `<Video 1>`, and `<Audio 1>`. State one explicit role for each source—identity, wardrobe, style, camera movement, performance timing, ambience, or voice—and resolve conflicts before generation.

How long can a MiniMax H3 workflow generate?

The released H3 specification supports 4–15 second output at 24 FPS. Local ComfyUI duration controls may snap to the model’s compatible frame grid. Longer or larger clips increase compute and memory, so validate the shot with a shorter preview before scaling.

Does MiniMax H3 generate audio in the same workflow?

Yes. H3 jointly generates video and 32 kHz stereo audio rather than adding a stock track afterward. Direct dialogue, effects, music, distance, and timing in the prompt, then review pronunciation, synchronization, clipping, unwanted sound, and usage rights before publishing.