MiniMax H3 Prompt Guide: Better Prompts for Every Video Workflow
Learn how to write MiniMax H3 prompts for text-to-video, image animation, end-frame control, reference media, camera movement, dialogue, and sound.

A good MiniMax H3 prompt reads less like an image caption and more like direction for a short scene. Name the subject, describe what visibly happens, place the action in a setting, direct the camera, and tell the model how the shot should sound and end. If you attach an image, video, or audio reference, give that file one clear job.
The most useful starting formula is:
Subject + visible action + setting + timing + camera + lighting/style + sound + ending state
That order is not a magic incantation. It is a checklist that removes ambiguity. MiniMax's official H3 prompt documentation likewise separates the audiovisual timeline, environmental sound, and background music, while its model card describes output with native stereo audio and clips up to 15 seconds. On Videoify's MiniMax H3 generator, you can apply the same ideas in a plain-language prompt for text-to-video, image-to-video, or reference-to-video.
This guide provides original starting templates rather than claims about guaranteed output. AI video generation varies from run to run, so treat each prompt as a controlled first draft.
Choose the workflow before writing the prompt
The same scene needs a different prompt depending on what the model already knows from your inputs. A text-only request must establish the entire shot. An image-to-video prompt should concentrate on movement instead of describing the uploaded frame back to the model.
| Videoify workflow | What the input already provides | What the prompt should emphasize |
|---|---|---|
| Text to video | Nothing visual | Subject, appearance, setting, action, camera, light, sound, and ending |
| Image to video | A required opening frame; optional end frame | Motion, pacing, camera behavior, details that must remain stable, and the destination |
| Reference to video | Images, video, or audio used as guidance | The role of each file, the new scene to create, what to preserve, and what may change |
This distinction prevents one of the most common prompt problems: spending half the prompt repeating information already visible in a source image while leaving the motion underspecified.
The MiniMax H3 prompt structure, one decision at a time
1. Start with one shot objective
Before adding cinematic language, decide what the shot needs to accomplish. For example:
-
Reveal a perfume bottle and finish on a clean product frame.
-
Animate a portrait with natural breathing and a slow turn toward the window.
-
Transfer the movement of a reference dancer to a different character.
-
Move from an approved closed-package frame to an approved open-package frame.
The objective is an editing decision. It tells you which details belong in the prompt and which ones are noise.
2. Describe visible action in playback order
“A tense woman in an office” describes a state. “She checks her watch, tightens her grip on a folder, and looks toward the closed meeting-room door” gives the model events it can place on a timeline.
Write actions in the order the viewer should see them. For a short clip, one clear action with a beginning and an end is usually easier to control than a chain of unrelated events.
3. Use camera terms precisely
Camera vocabulary changes the viewpoint, so choose the term that matches the movement you mean:
| Prompt term | What it asks the camera to do |
|---|---|
| Push in / pull out | Move the camera physically closer to or farther from the subject |
| Zoom in / zoom out | Change focal length while the camera stays in place |
| Pan left / right | Rotate the camera horizontally from a fixed position |
| Truck left / right | Move the whole camera sideways |
| Tilt up / down | Rotate the camera vertically from a fixed position |
| Pedestal up / down | Raise or lower the whole camera |
| Tracking shot | Follow a moving subject |
| Arc shot | Move around the subject on a curved path |
| Static shot | Hold both position and framing steady |
MiniMax's official base prompt guide recommends expressing camera movement as part of the shot rather than stacking camera labels at the end. In practice, “the camera slowly pushes toward the watch in her hand” is clearer than “cinematic, dolly, zoom, orbit, dynamic camera.”
Start with one main camera move. Add a second only when the shot genuinely needs it and there is enough time for both.

Camera rotation, camera translation, and lens changes create different motion even when the subject stays still. Conceptual illustration.
4. Treat sound as part of the scene
MiniMax H3 generates picture and audio together, according to the official model card. That makes silence about sound a creative choice: the model still has to decide what the world sounds like.
Separate three kinds of audio in your head, even if you write them in one paragraph:
-
Dialogue: the exact words, the speaker, delivery, and timing.
-
Diegetic sound: sound inside the scene, such as footsteps, fabric, rain, machinery, or a radio.
-
Background score: music meant for the audience rather than the characters.
Instead of “dramatic audio,” try “distant traffic under steady rain; the doorbell rings when she enters; no background music.” Concrete sources and timing give the model more useful direction than mood alone.
5. Finish with the state you want to keep
The ending is easy to overlook, but it often determines whether a clip is useful in an edit. Ask for a stable final composition: the product label faces camera, the character stops in profile, the door closes, or the camera settles on a wide view.
When you provide an end frame, describe the visible path toward it. Do not simply tell the model to “transition smoothly.” Say which object moves, how the pose changes, how the camera reframes, and what should match when the clip lands.
Copy-ready MiniMax H3 prompt templates
Replace the bracketed text with your own details. Keep only the fields that affect the shot.
Text-to-video prompt
Use text-to-video when the model must build the complete scene from language.
[Visual style and framing]. [Subject with stable identifying details] is in [setting].
[Subject] [performs one visible action in chronological order].
The camera [one camera movement, speed, and destination].
[Lighting and color details that matter].
Sound: [ambience and physical sounds]. [Speaker] says, "[short line]," in [delivery].
End with [final pose or composition]. Preserve [critical details].
Example:
Live-action product film, medium close-up. A cobalt-blue perfume bottle stands on wet black stone in a dark studio. A thin ribbon of mist moves behind the bottle as droplets slide down the glass. The camera makes a slow half-arc from left to right and settles directly in front of the label. Narrow white highlights trace the bottle edges; the background stays charcoal black. Sound: soft water drops and a low room tone, no music. End on a stable, centered product frame with the label facing the camera and the bottle shape unchanged.
An unbranded cobalt perfume bottle reveal that moves from an angled view to a stable front-facing product frame, with synchronized water sounds and no music.
Image-to-video motion prompt
Your opening image already defines appearance, composition, and much of the lighting. Prompt the change.
Begin from the uploaded first frame. [Subject] [starts with subtle motion], then [main action].
The camera [movement] while [specific identity, clothing, product, or background details] remain unchanged.
Sound: [scene audio]. End with [final visible state].
Example:
Begin from the uploaded portrait. The woman takes a natural breath, blinks once, then turns her eyes toward the rain outside the train window. Reflections slide slowly across the glass while the camera trucks right with a small, steady movement. Keep her facial features, short black hair, red scarf, seat position, and the carriage layout unchanged. Sound: soft wheel rhythm, ventilation hum, and light rain on glass. End with her gaze held on the passing city lights.
Avoid redescribing every pixel in the image. Repeat only the details that are likely to drift or that the shot depends on.
First-and-last-frame prompt
With two frame images, the prompt's job is to explain the motion between them. MiniMax's official guide favors a continuous path for this workflow, especially when no cut is requested.
Start exactly from the first frame and end on the composition in the final frame.
[Object or subject] moves from [opening state] to [ending state] by [observable intermediate actions].
The camera [continuous camera behavior]. Preserve [details shared by both frames].
Sound: [sounds synchronized to the transition].
Example:
Start exactly from the closed-box first frame and end on the open-box final frame. The paper seal lifts at one corner, peels back across the lid, and falls beside the package. The lid then rises on its hinge as the camera slowly pushes in, revealing the watch without changing its position. Preserve the box proportions, cream paper texture, gold logo, tabletop, and warm side lighting. Sound: a soft paper peel, a small hinge click, and quiet studio ambience. Settle on the final composition with the watch fully visible and the camera still.
If the start and end images differ in too many unrelated ways—subject, perspective, background, lighting, and layout—the model has to invent several transitions at once. Use more compatible endpoints or generate the change as separate shots.
Reference-to-video prompt
The reference workflow is for guidance, not deterministic frame-by-frame editing. On Videoify, switch to Reference to Video before adding image, video, or audio files, then state what each source should control.
Use Image 1 for [character or product identity].
Use Image 2 for [wardrobe, location, composition, or style].
Use Video 1 only for [movement, timing, or camera path].
Use Audio 1 for [rhythm, vocal quality, or pacing].
Create a new [duration]-second shot in which [subject] [action] in [setting].
The camera [movement]. Preserve [identity details] from [reference].
Do not copy [an unwanted element from a reference].
End with [final state].
Example:
Use Image 1 for the courier's face and short silver hair. Use Image 2 for the orange raincoat and black messenger bag. Use Video 1 only for the runner's brisk stride and the low side-tracking camera. Use Audio 1 for the percussion rhythm, not for an exact voice copy.
Create a new night-market shot in which the courier moves quickly through a narrow aisle, dodges one cart, and looks over her shoulder once. Keep her face, hair, coat, and bag consistent. Neon reflections move across the wet pavement as the camera tracks beside her at natural speed. End as she stops beneath a blue shop awning. Do not carry the original runner's clothing or background into the new scene.
More references are not automatically better. Use the smallest consistent set that communicates the subject, movement, and environment. Conflicting faces, outfits, camera angles, or lighting cues make the model's job less clear.
Dialogue and native-audio prompt
Keep spoken material realistic for a 4–15 second clip. State who speaks, quote the exact line, and describe the delivery.
The baker looks toward the customer and says, "First batch of the morning," in a warm, slightly raspy voice. She begins speaking after the shutters open and finishes before the camera cuts to the bread. Her lips move with the line. Sound: wooden shutters scraping, one tray placed on the counter, and quiet street ambience. No background music.
The official H3 model card lists stable dialogue support for Arabic, Chinese, English, French, German, Italian, Japanese, Korean, Portuguese, Russian, and Spanish, with varying support beyond those languages. Videoify does not publish a separate validated language list for its hosted workflow, so pronunciation, timing, exact wording, and lip synchronization should still be reviewed after every run.
How to direct timing and multiple shots
For a single shot, ordinary sequence words are often enough: “first,” “then,” “as,” and “finally.” Add explicit timing when an event must happen near a particular point.
0–3 seconds: the cyclist waits beside the closed gate as the camera holds a wide shot.
3–6 seconds: the gate rises and she starts forward at natural speed.
6–8 seconds: the camera tracks beside her and settles as she exits into sunlight.
MiniMax H3 supports multi-shot generation, but more cuts create more opportunities for identity and spatial continuity to drift. If a clip has one purpose, keep it as one shot. If it needs a cut, use the cut to reveal new information and place it at a clear time. For a longer story, generate separate clips with the same strong reference set and assemble them in an editor; Videoify produces individual clips rather than a full editing timeline.
Fix common MiniMax H3 prompt failures
| What you see | Likely prompt problem | What to change |
|---|---|---|
| The clip feels like slow motion | The action has no clear pace, or words such as “dreamy” and “graceful” imply slowness | Ask for “natural speed,” “real-time movement,” or a brisk pace; give the action a clear start and finish |
| The subject changes appearance | Identity details are vague or the reference set conflicts | Use fewer, consistent images and repeat the features that must stay stable |
| The camera ignores the direction | Several movements compete in one short shot | Keep one primary movement and name its speed and destination |
| The end frame arrives as a jump | The prompt names the endpoints but not the path | Describe the intermediate pose, object, camera, and lighting changes |
| Dialogue is garbled or late | The line is too long or the speaker/timing is unclear | Shorten the line, identify the speaker, and say when speech begins and ends |
| The result is visually crowded | The prompt contains several subjects, actions, styles, and props with equal priority | Return to one shot objective and remove details that do not change the result |
| Unwanted text appears | A sign, label, or interface element is ambiguous | Quote only the text that must appear; otherwise ask for a clean surface and add precise typography later in an editor |
| A reference dominates the wrong feature | The file's role is unstated | Assign each reference one job and say which elements should not transfer |
Change one variable per retry. If you rewrite the action, camera, lighting, and reference set simultaneously, a better result will not tell you which change helped.
A controlled testing loop on Videoify
Prompt iteration is part of the workflow. Start with the simplest useful version so each retry tells you something specific.
-
Start with one subject, one action, one camera movement, and a clear ending.
-
Run a short draft to check motion, composition, and timing.
-
Watch the whole clip with sound. Note the first visible failure, not every possible improvement.
-
Change one instruction or one reference and generate again.
-
Once the direction works, generate the version you intend to keep.
Keep a short record of the workflow, prompt, duration, references, and the single change made between versions. That makes it easier to repeat a successful setup or identify the instruction that caused a new problem.
Final prompt checklist
Before generating, make sure you can answer these questions from the prompt:
-
What is the shot trying to accomplish?
-
Who or what is the main subject?
-
What visibly changes, and in what order?
-
How does the camera move—or should it stay still?
-
What should the scene sound like?
-
Which details must remain consistent?
-
What does the final frame look like?
-
If references are attached, does each one have a named role?
-
Can the action and dialogue plausibly fit within the selected duration?
If the answer is yes, the prompt is probably detailed enough. Length by itself is not the goal. A focused paragraph can give MiniMax H3 better direction than a page of adjectives.
When your prompt is ready, choose the matching workflow in the MiniMax H3 workspace, run a short draft, and revise from what the clip actually does.