T2VA · I2VA · FL2VA · L2VA · camera · dialogue · sound
MiniMax H3 Prompt GuideWrite Timed Shots, Clear Camera Moves, and Better Audio
A useful H3 prompt is a production brief, not a pile of visual adjectives. Choose the input task, describe observable change over time, direct camera and cuts, preserve exact dialogue, and separate the scene soundscape from non-diegetic music.
Independent service. Not affiliated with, endorsed by, or sponsored by MiniMax or any model owner.
MiniMax H3 prompt guide / official format checked August 8, 2026
Write what the audience can see and hear over time
A MiniMax H3 prompt should identify the task, establish the opening state, sequence visible action and camera changes, preserve required speech, and finish with an explicit sound plan. The official writing framework separates the integrated multimodal description from the overall soundscape and non-diegetic music. This page translates that framework into a practical browser workflow; it does not claim that this service runs the official Prompt Rewriter automatically.
Match the prompt timeline to T2VA, I2VA, FL2VA, or L2VA
T2VA begins with no visual anchor, so the prompt must establish subject, environment, composition, action, camera, lighting, pace, and sound. I2VA starts from a supplied first frame: acknowledge that visible state, then describe action onset and continuous development instead of redrawing the image in words. The frame owns the opening geometry; the prompt owns what changes after it.

FL2VA uses both first and last frames. State the opening condition, name the intermediate physical and camera changes, and land naturally on the supplied destination. L2VA works backward from a required last frame: propose a plausible earlier state, describe convergence, and reserve the final beat for the exact landing. Do not demand an ending that contradicts the uploaded composition or a transition too complex for the selected duration.
Build one integrated description as a timed shot plan
Start with a compact setup: subject, location, time, composition, lighting, and the continuity details that cannot drift. Continue with observable verbs in chronological order. For a ten-second product shot, for example: the sealed case rests on wet stone; at two seconds the latch releases; from three to six seconds the camera pushes in as vapor rolls outward; by nine seconds the product reaches the supplied hero angle. Concrete state changes are easier to evaluate than phrases such as ‘epic cinematic energy.’

If the piece needs multiple shots, declare Shot 1 and then use strictly increasing cut times from Shot 2 onward. A cut at 00:04 must be followed by a later cut, never another 00:04 or an earlier timestamp. Keep each segment feasible: one dominant action, one camera intention, and a clear continuity bridge. A sequence with six cuts, three costume changes, two locations, and a complex stunt in eight seconds is a scheduling contradiction, not a detailed prompt.
Direct camera motion with type, amplitude, and speed
Name what the camera does, how far it moves, and how quickly. ‘Slow, shallow dolly-in from a medium shot to a chest-up portrait’ is testable; ‘dynamic camera’ is not. Distinguish camera movement from subject movement: the runner accelerates left to right while the camera tracks parallel at waist height. Add lens feeling, depth of field, or handheld texture only when it changes the intended image, and avoid stacking mutually exclusive directions such as locked-off, handheld, and rapid orbit in one beat.

For cuts, explain the visual reason for the transition. Cut from a wide environment reveal to a macro material detail when the sound cue lands; then return to a stable hero frame. Preserve screen direction, wardrobe, lighting logic, and object state across shots unless a deliberate change is written. When references are attached, give each source one job—identity, motion rhythm, camera language, environment, or voice—and state which property wins if two sources disagree.
Keep dialogue, soundscape, and music in separate lanes
Use stable speaker IDs from beginning to end, especially when two people speak. Write exact lines in quotation marks and do not paraphrase them elsewhere. Place the line after the action that motivates it, then specify delivery briefly: Speaker A whispers, ‘Do not turn around,’ while keeping eye contact with the doorway. Avoid assigning the same line to multiple labels or asking a speaker to deliver more words than the available seconds reasonably allow.

Describe the overall soundscape as sounds that belong to the scene: room tone, footsteps, fabric, machinery, weather, impacts, crowd distance, and spatial placement. Describe non-diegetic music separately: instrumentation, tempo, emotional arc, start or stop point, and whether it ducks beneath dialogue. Do not ask for silence and loud music at the same time. Review the result for intelligibility, lip sync, pronunciation, clipping, unwanted score, stereo balance, and whether every important sound arrives with the visible event.
Choose the prompt pattern that matches the input
Open text concept. Use T2VA: establish the whole visual world, then sequence subject action, camera motion, cuts, dialogue, effects, and music through the selected duration.
Keyframe-directed transition. Use I2VA, FL2VA, or L2VA: treat supplied frames as geometric anchors and spend the words on plausible temporal change between them.
Reference-led performance. Assign image identity, video rhythm, and audio or voice roles explicitly; remove references that compete for the same property.
Continue through the MiniMax H3 production guides
Give every sentence one production job.
A reliable prompt separates the visual timeline, camera plan, exact speech, scene sound, and music. The task mode determines where that timeline begins and ends.
Anchor, sequence, then listen.
Choose T2VA, I2VA, FL2VA, or L2VA. State the opening condition, required continuity, and any supplied frame or reference role.
Write observable actions and camera changes in order. Use strictly increasing cut times and keep each shot feasible within the duration.
Preserve exact dialogue with stable speaker IDs, then specify scene sound and non-diegetic music separately before generating.
Turn the shot plan into an H3 generation.
Choose text, keyframes, or multimodal references above, paste a timed visual brief, and direct dialogue, effects, ambience, and music in the same run.
Test prompts with repeatable settings
Use credits to compare candidates at the same mode, duration, resolution, and reference set before changing the next high-value variable.
Standard
Billed $238.80 yearly
- Commercial Usage RightsYearly Only
- Up to 624 videos
- Video models low to $0.11/s
- Up to 9,360 images
- All AI Models
- Priority Support
- Priority Processing Speed
- Remove Watermark
- Free Video UpscalerYearly Only
- Free Frame InterpolationYearly Only
- More Free AI ToolsYearly Only
Business
Billed $418.80 yearly
- Commercial Usage RightsYearly Only
- Up to 1,360 videos
- Video models low to $0.09/s
- Up to 20,400 images
- All AI Models
- Priority Support
- Dedicated Support
- Priority Processing Speed
- Remove Watermark
- Free Video UpscalerYearly Only
- Free Frame InterpolationYearly Only
- More Free AI ToolsYearly Only
Enterprise
Billed $898.80 yearly
- Commercial Usage RightsYearly Only
- Up to 3,920 videos
- Video models low to $0.06/s
- Up to 58,800 images
- All AI Models
- Priority Support
- Dedicated Support
- Priority Processing Speed
- Remove Watermark
- Free Video UpscalerYearly Only
- Free Frame InterpolationYearly Only
- More Free AI ToolsYearly Only
Included models
Credits activate instantly for MiniMax H3 generation workflows.

If a direction cannot be seen, heard, or timed, rewrite it.
Replace vague praise with a visible state change, a measurable camera move, an exact line, a synchronized sound, or a clearly assigned reference role.
Test a Shot-Ready PromptMiniMax H3 Prompt Guide FAQ
How do I write a MiniMax H3 prompt?
Write it as a timed production brief: choose the task mode, establish the opening state, describe observable action and camera movement in order, preserve exact dialogue with speaker IDs, then separate scene sound from non-diegetic music. Generate several candidates with the same settings before revising one variable.
How do T2VA, I2VA, FL2VA, and L2VA prompts differ?
T2VA must invent the full scene. I2VA continues from a supplied first frame. FL2VA describes a plausible transition between supplied first and last frames. L2VA constructs an earlier state that converges on the supplied ending. In every mode, the prompt should describe temporal change rather than contradicting the visual anchor.
How should I time cuts in a MiniMax H3 prompt?
Start with Shot 1, then give every later cut a strictly increasing timestamp. Keep one dominant action and camera intention per segment, preserve continuity across the edit, and leave enough seconds for speech and physical movement. Too many cuts in a short duration create conflicting instructions rather than useful detail.
How should I describe camera movement for MiniMax H3?
Specify motion type, amplitude, and speed: for example, a slow shallow dolly-in from medium shot to chest-up. Separate camera movement from subject movement, and add lens or handheld characteristics only when they matter. Avoid combining incompatible directions such as locked-off, rapid orbit, and handheld shake in one beat.
How do I prompt dialogue and native audio?
Use stable speaker IDs and quote the exact words once, positioned after the action that motivates them. Describe room tone, effects, weather, footsteps, and spatial sound as the overall soundscape. Describe music separately with instrumentation, tempo, timing, and ducking so it does not obscure dialogue.
Does this site use the official MiniMax H3 Prompt Rewriter?
This page does not claim that the browser generator runs the official Prompt Rewriter. It explains a practical writing method based on the published task structure. What you enter is the creative brief you should review, save, and compare; any provider-side processing is not presented here as a user-controlled rewriter feature.