MiniMax H3 AI Video Generator

Create a MiniMax H3 video from text, first or last frames, or mixed image, video, and audio references with visible duration, resolution, framing, and credit controls.

Model
MiniMax H3
Inputs
Text · Frames · Mixed refs
Output
4–15s · 768p / 2K

Loading generator…

What Is MiniMax H3?

Editorial desk with storyboard frames, reference media, and audio cues
ZMS editorial planning artwork for organizing storyboard beats, frame anchors, reference clips, and audio cues before submission. It is not a MiniMax H3 output, task record, result preview, or quality proof.

MiniMax H3 is a specific multimodal video model, not a MiniMax video-family label. In the ZMS AI workspace, prepare text-to-video, first- or last-frame, and mixed-reference tasks with 4–15 second duration and 768p or 2K resolution. Its upstream identity covers text, image, video, and audio context plus native stereo sound; this page documents the narrower ZMS task contract.

Visible controls and Generate video connect to the current task flow, but do not prove signed-in completion, playable media, or final debit. Until that end-to-end check is complete, this page makes no performance claim and never treats planning artwork as model output.

Choose a MiniMax H3 Input Mode

Choose the interface that matches your source material. For broader model selection, compare AI Video Generator routes before uploading.

Illustrated planning routes for text, frame anchors, and mixed references
This ZMS planning illustration separates a written brief, anchored frames, and mixed-media roles for comparing the three interfaces. It is not an upload capture, generated frame, or MiniMax H3 result.
ModePrepareInterface boundary
Text to videoA timeline prompt covering subject, action, camera, setting, and sound.No upload. Choose one of six text ratios before generating.
First / last frameA first frame, a last frame, or both, plus the intended motion.At least one frame is required. This interface does not send a ratio field.
Mixed referencesUp to 5 images, 3 videos, and 3 audio files, each with a named role.Audio cannot be the only reference; include at least one image or video.

Set Duration, Resolution, and Framing

Settings change with the selected input mode. The live selector is the source of truth; these cards explain where each choice applies.

4–15 seconds

Set an integer duration from 4 through 15 seconds. Use the shortest span that still lets the action, camera move, and sound cue read.

768p or 2K

Choose 768p for a lower estimate or 2K for the higher-resolution path. These are selectable sizes, not measured quality claims.

Six text ratios

Text mode offers 21:9, 16:9, 4:3, 1:1, 3:4, and 9:16. Match the frame to its display rather than asking the prompt to override it.

Mode-aware framing

Reference mode can use adaptive framing. First/last-frame mode sends no ratio because supplied anchors establish the frame.

Estimate MiniMax H3 Credits Before You Generate

The workspace estimates from the shared pricing helper. Credits belong to your ZMS AI plan; the service response remains authoritative for final debit.

Task caseVisible rateEstimate
768p output20 credits per secondNo reference video: output seconds × 20.
2K output35 credits per secondNo reference video: output seconds × 35.
With reference videoSelected resolution rate(output seconds + total reference-video seconds) × rate.

Try Four MiniMax H3 Prompt Recipes

Each original recipe maps to a real input path and fills safe text and settings without starting a task. The mixed-reference recipe complements the mixed-reference workflow with H3-specific limits and a different brief.

MiniMax H3 starters

4 starter recipes
Text to Video

Studio product reveal

Plan one reveal, camera move, and sound beat.

Duration
8 seconds
Frame
16:9
Resolution
768p
Audio
Ambient cue
Suggested prompt

0–2s: a ceramic speaker rests on a blue plinth under rim light. 2–6s: orbit clockwise as fabric panels open. 6–8s: settle on the logo side under warm light. Sound: one synth pulse, then a quiet click; no dialogue.

Review checkpoint

Check silhouette, orbit direction, proportions, and sound timing.

First Frame

Continue a market walk

Continue one opening frame with restrained handheld motion.

Duration
7 seconds
Frame
From frame
Resolution
768p
Audio
Location ambience
Suggested prompt

Continue from the first frame. The subject walks toward the fruit stall, turns left, and reaches for an orange. Follow at shoulder height with restrained drift; preserve clothing, face, layout, and late-afternoon light. Add market ambience and footsteps; no speech.

Review checkpoint

Inspect identity, hand motion, camera height, and background geometry.

First + Last Frames

Bridge day into night

Bridge two supplied anchors with one motivated transition.

Duration
10 seconds
Frame
From frames
Resolution
2K
Audio
Environmental transition
Suggested prompt

Begin at the daylight rooftop frame and reach the night frame. Push toward the railing as clouds accelerate, windows illuminate, and the sky turns deep blue. Keep skyline, lens height, and railing stable. Sound: wind softens as city ambience rises.

Review checkpoint

Check the last composition for jumps, duplicates, or abrupt audio.

Mixed References

Editorial travel vignette

Assign separate visual, motion, and sound roles.

Duration
12 seconds
Frame
Adaptive
Resolution
2K
Audio
Uploaded wave texture
Suggested prompt

Use image 1 for the traveler’s coat, image 2 for the station, video 1 for lateral tracking, and audio 1 for waves. The traveler crosses the platform, pauses as a train enters, then looks toward the sea. Preserve identity and layout; do not copy people from the motion reference. No dialogue.

Review checkpoint

Include a visual reference and keep motion, identity, and audio distinct.

How to Use the MiniMax H3 Video Generator

Build one deliberate task from inputs to review instead of treating every available file as a useful reference. Choose a mode first, give each upload a named purpose, and keep the timeline short enough to evaluate. Before uploading personal or third-party media, check the ZMS AI terms and confirm that you have the necessary rights.

Editorial board mapping frame, reference, audio, and review roles
This editorial planning board maps frame anchors, visual references, motion references, audio roles, and final review checks before a task is submitted. It is explanatory ZMS artwork, not a MiniMax H3 video frame or completed generation record.
  1. 01

    Choose the matching input path

    Use Text when the brief stands alone, First/last frame for one or two visual anchors, or Reference when multiple media roles must guide the sequence.

  2. 02

    Write a timed, role-aware prompt

    Describe subject, action, camera, setting, and sound in chronological beats. Set duration and resolution, then choose a ratio only when that interface exposes one.

  3. 03

    Check uploads and the estimate

    Confirm reference labels, the 5/3/3 limits, at least one visual reference in mixed mode, and the added input-video seconds shown in the credit estimate.

  4. 04

    Generate and read the real state

    Select Generate video, complete sign-in if requested, and follow the actual validation, task, error, and returned-media state. Do not treat an empty result panel as an example gallery.

Compact timeline pattern

Create a restrained ten-second sequence with one continuous visual idea. 0–3s: establish the main subject in a readable environment, hold the first composition long enough to identify the subject, and introduce quiet location ambience. 3–7s: perform one clearly described action while the camera makes one motivated move; preserve identity, wardrobe, object proportions, lighting direction, and background geometry. 7–10s: slow the action and settle on the intended final composition without an abrupt cut or last-second pose change. If references are attached, use each only for its named role and do not copy unrelated people or objects. Keep motion physically coherent and avoid flicker, duplicated limbs, warped text, or sudden changes in lens height. Name ambience, effects, music, or dialogue only when each sound has a clear role, timing, and stopping point.

  • Every uploaded reference has one named role.
  • Duration, resolution, and ratio match the selected mode.
  • The final frame and audio beat have an explicit stopping point.

Know What Is Connected and What Still Needs Review

This ledger separates the visible implementation from evidence that still requires a signed-in task and a recorded response. It should be updated when an acceptance record changes, not converted into a permanent marketing claim or treated as proof of output quality.

StateWhat is true nowHow to use it
Visible workspaceThree modes, mode-aware fields, upload limits, a shared credit estimator, and Generate video are connected to the current H3 task contract.Configure a real task, check the current mode summary, and rely on visible validation before submission.
Not independently verifiedSigned-in completion, status polling, returned playable media, and the final account debit are not yet recorded together as a ZMS acceptance result.Do not infer speed, quality, reliability, delivery success, or a completed charge from page availability alone.
Publication reviewA returned task still needs rights, identity, continuity, readable text, motion, sound, and error-state review before publication.Keep only media you are authorized to use, retain task provenance, and review the actual response before publishing it.

MiniMax H3 Generator FAQ

These answers cover the current H3 workspace contract and its evidence boundary. When you are ready, open the MiniMax H3 workspace and review the estimate before submitting.

Can I generate with MiniMax H3 on ZMS AI?

The page exposes a real Generate video action connected to the current create and status-check flow. A signed-in task still needs end-to-end acceptance for completed media and final debit, so availability is not presented as an independent performance test.

Which MiniMax H3 input mode should I choose?

Choose Text for a prompt-only task, First/last frame when one or two anchor images should define the visual path, and Reference when images, videos, or audio need separate guiding roles.

How many images, videos, and audio files can I add?

Reference mode accepts up to five images, three videos, and three audio files. First/last-frame mode instead provides two dedicated image slots and requires at least one of them.

Can audio be the only reference?

No. A mixed-reference task must contain at least one visual reference—an image or a video. Audio can guide ambience, effects, music, or another sound role alongside that visual material.

Which durations, resolutions, and ratios are available?

All three modes support 4–15 seconds and 768p or 2K. Text offers 21:9, 16:9, 4:3, 1:1, 3:4, and 9:16; Reference can use adaptive; First/last frame sends no ratio field.

How are MiniMax H3 credits estimated?

768p uses 20 credits per output second and 2K uses 35. If reference videos are attached, their total seconds are added to output duration before multiplying by the selected rate. Final debit remains service-authoritative.

Are the planning images MiniMax H3 outputs?

No. All three images are page-specific ZMS editorial artwork that explains input planning. They are not generated results, task records, interface captures, or evidence of MiniMax H3 quality.