4–15 seconds
Set an integer duration from 4 through 15 seconds. Use the shortest span that still lets the action, camera move, and sound cue read.
Create a MiniMax H3 video from text, first or last frames, or mixed image, video, and audio references with visible duration, resolution, framing, and credit controls.
Loading generator…
Model overview

MiniMax H3 is a specific multimodal video model, not a MiniMax video-family label. In the ZMS AI workspace, prepare text-to-video, first- or last-frame, and mixed-reference tasks with 4–15 second duration and 768p or 2K resolution. Its upstream identity covers text, image, video, and audio context plus native stereo sound; this page documents the narrower ZMS task contract.
Visible controls and Generate video connect to the current task flow, but do not prove signed-in completion, playable media, or final debit. Until that end-to-end check is complete, this page makes no performance claim and never treats planning artwork as model output.
Input planning
Choose the interface that matches your source material. For broader model selection, compare AI Video Generator routes before uploading.

| Mode | Prepare | Interface boundary |
|---|---|---|
| Text to video | A timeline prompt covering subject, action, camera, setting, and sound. | No upload. Choose one of six text ratios before generating. |
| First / last frame | A first frame, a last frame, or both, plus the intended motion. | At least one frame is required. This interface does not send a ratio field. |
| Mixed references | Up to 5 images, 3 videos, and 3 audio files, each with a named role. | Audio cannot be the only reference; include at least one image or video. |
Output controls
Settings change with the selected input mode. The live selector is the source of truth; these cards explain where each choice applies.
Set an integer duration from 4 through 15 seconds. Use the shortest span that still lets the action, camera move, and sound cue read.
Choose 768p for a lower estimate or 2K for the higher-resolution path. These are selectable sizes, not measured quality claims.
Text mode offers 21:9, 16:9, 4:3, 1:1, 3:4, and 9:16. Match the frame to its display rather than asking the prompt to override it.
Reference mode can use adaptive framing. First/last-frame mode sends no ratio because supplied anchors establish the frame.
Credit estimate
The workspace estimates from the shared pricing helper. Credits belong to your ZMS AI plan; the service response remains authoritative for final debit.
| Task case | Visible rate | Estimate |
|---|---|---|
| 768p output | 20 credits per second | No reference video: output seconds × 20. |
| 2K output | 35 credits per second | No reference video: output seconds × 35. |
| With reference video | Selected resolution rate | (output seconds + total reference-video seconds) × rate. |
Example: 8 seconds at 768p estimates 160 credits. Add 6 reference-video seconds and the estimate becomes 14 × 20 = 280. Images remain capped at five; no sixth-image overage exists.
Compare credit plansPrompt starters
Each original recipe maps to a real input path and fills safe text and settings without starting a task. The mixed-reference recipe complements the mixed-reference workflow with H3-specific limits and a different brief.
MiniMax H3 starters
4 starter recipesPlan one reveal, camera move, and sound beat.
0–2s: a ceramic speaker rests on a blue plinth under rim light. 2–6s: orbit clockwise as fabric panels open. 6–8s: settle on the logo side under warm light. Sound: one synth pulse, then a quiet click; no dialogue.
Check silhouette, orbit direction, proportions, and sound timing.
Continue one opening frame with restrained handheld motion.
Continue from the first frame. The subject walks toward the fruit stall, turns left, and reaches for an orange. Follow at shoulder height with restrained drift; preserve clothing, face, layout, and late-afternoon light. Add market ambience and footsteps; no speech.
Inspect identity, hand motion, camera height, and background geometry.
Bridge two supplied anchors with one motivated transition.
Begin at the daylight rooftop frame and reach the night frame. Push toward the railing as clouds accelerate, windows illuminate, and the sky turns deep blue. Keep skyline, lens height, and railing stable. Sound: wind softens as city ambience rises.
Check the last composition for jumps, duplicates, or abrupt audio.
Assign separate visual, motion, and sound roles.
Use image 1 for the traveler’s coat, image 2 for the station, video 1 for lateral tracking, and audio 1 for waves. The traveler crosses the platform, pauses as a train enters, then looks toward the sea. Preserve identity and layout; do not copy people from the motion reference. No dialogue.
Include a visual reference and keep motion, identity, and audio distinct.
Four-step workflow
Build one deliberate task from inputs to review instead of treating every available file as a useful reference. Choose a mode first, give each upload a named purpose, and keep the timeline short enough to evaluate. Before uploading personal or third-party media, check the ZMS AI terms and confirm that you have the necessary rights.

Use Text when the brief stands alone, First/last frame for one or two visual anchors, or Reference when multiple media roles must guide the sequence.
Describe subject, action, camera, setting, and sound in chronological beats. Set duration and resolution, then choose a ratio only when that interface exposes one.
Confirm reference labels, the 5/3/3 limits, at least one visual reference in mixed mode, and the added input-video seconds shown in the credit estimate.
Select Generate video, complete sign-in if requested, and follow the actual validation, task, error, and returned-media state. Do not treat an empty result panel as an example gallery.
Compact timeline pattern
Create a restrained ten-second sequence with one continuous visual idea. 0–3s: establish the main subject in a readable environment, hold the first composition long enough to identify the subject, and introduce quiet location ambience. 3–7s: perform one clearly described action while the camera makes one motivated move; preserve identity, wardrobe, object proportions, lighting direction, and background geometry. 7–10s: slow the action and settle on the intended final composition without an abrupt cut or last-second pose change. If references are attached, use each only for its named role and do not copy unrelated people or objects. Keep motion physically coherent and avoid flicker, duplicated limbs, warped text, or sudden changes in lens height. Name ambience, effects, music, or dialogue only when each sound has a clear role, timing, and stopping point.
Verification boundary
This ledger separates the visible implementation from evidence that still requires a signed-in task and a recorded response. It should be updated when an acceptance record changes, not converted into a permanent marketing claim or treated as proof of output quality.
| State | What is true now | How to use it |
|---|---|---|
| Visible workspace | Three modes, mode-aware fields, upload limits, a shared credit estimator, and Generate video are connected to the current H3 task contract. | Configure a real task, check the current mode summary, and rely on visible validation before submission. |
| Not independently verified | Signed-in completion, status polling, returned playable media, and the final account debit are not yet recorded together as a ZMS acceptance result. | Do not infer speed, quality, reliability, delivery success, or a completed charge from page availability alone. |
| Publication review | A returned task still needs rights, identity, continuity, readable text, motion, sound, and error-state review before publication. | Keep only media you are authorized to use, retain task provenance, and review the actual response before publishing it. |
FAQ
These answers cover the current H3 workspace contract and its evidence boundary. When you are ready, open the MiniMax H3 workspace and review the estimate before submitting.
The page exposes a real Generate video action connected to the current create and status-check flow. A signed-in task still needs end-to-end acceptance for completed media and final debit, so availability is not presented as an independent performance test.
Choose Text for a prompt-only task, First/last frame when one or two anchor images should define the visual path, and Reference when images, videos, or audio need separate guiding roles.
Reference mode accepts up to five images, three videos, and three audio files. First/last-frame mode instead provides two dedicated image slots and requires at least one of them.
No. A mixed-reference task must contain at least one visual reference—an image or a video. Audio can guide ambience, effects, music, or another sound role alongside that visual material.
All three modes support 4–15 seconds and 768p or 2K. Text offers 21:9, 16:9, 4:3, 1:1, 3:4, and 9:16; Reference can use adaptive; First/last frame sends no ratio field.
768p uses 20 credits per output second and 2K uses 35. If reference videos are attached, their total seconds are added to output duration before multiplying by the selected rate. Final debit remains service-authoritative.
No. All three images are page-specific ZMS editorial artwork that explains input planning. They are not generated results, task records, interface captures, or evidence of MiniMax H3 quality.