Text to Video
Use text when composition is open. Define one subject, one action, one camera move, the environment, and an ending state; avoid competing events in the same short shot.
Build a Wan 3.0 video task from text, a required first frame with an optional last frame, or mixed references, with 2–30 second controls and credit estimates.
Loading generator…
Model workspace
The Wan 3.0 AI Video Generator on ZMS AI turns one shot idea into a structured task with its prompt, input mode, frame, duration, resolution, audio choice, and credit estimate in one workspace.

It supports text, a required first frame with an optional last frame, and mixed image, video, and audio references. Choose the smallest mode that protects the facts your shot cannot lose.
Use the Wan model family page if the version is undecided. ZMS has implemented the controls and adapter, but has not yet verified a complete signed-in generation, debit, and download path.
Choose an input path
Choose the mode by the anchor your shot needs. Extra media helps only when every file has a distinct job.
Use text when composition is open. Define one subject, one action, one camera move, the environment, and an ending state; avoid competing events in the same short shot.
Upload one required first frame; add one optional last frame only when the ending matters. Describe a plausible transition between anchors with compatible subject, light, geometry, and camera position.
Use mixed references when appearance, movement, or sound needs a separate source. Name each file's job and remove conflicts; the exact 10 / 5 / 5 limits are listed below.
Prompt-first starters
These are original preparation recipes, not prompts recovered from finished Wan 3.0 outputs. Choose Try prompt to move the full text into the workbench, then adjust the settings and review checkpoint for your own subject. Image recipes also require the clearly labeled reference input to be downloaded and uploaded by you.
Text to Video
4 starter recipesPractice a restrained product move in which the camera supplies motion while the object remains geometrically stable.
A single matte-black travel speaker stands centered on a pale stone plinth in a quiet daylight studio. Keep the speaker's rectangular proportions, grille pattern, controls, logo placement, and edges unchanged. The camera makes one slow 120-degree clockwise orbit at product height while a soft window reflection travels across the surface. No hands, no extra objects, no transformation, no cuts, no zoom. End on a clean three-quarter view with the full product inside frame.
Check the silhouette and grille before judging style; reject any orbit that bends the product, changes its details, or accelerates near the end.
Use one readable human action and a fixed camera to make facial, hand, and timing errors easier to identify.
Medium shot of an adult ceramic artist seated at a wooden workbench beside a north-facing window. The camera stays locked. The artist studies a small blue bowl, takes one calm breath, brushes a speck of clay from the rim with the right thumb, then looks toward the window. Natural overcast light brightens slightly across the shot. Preserve the same face, hands, clothing, bowl shape, and background layout. No dialogue, no camera movement, no additional action, no cut.
Review the face, thumb contact, bowl geometry, and gaze transition frame by frame; the beat should feel continuous rather than assembled from separate poses.
Build spatial clarity with one forward move, distinct depth layers, and a specified ending composition.
Early morning inside a narrow greenhouse aisle after rain. Begin behind out-of-focus fern leaves in the foreground, with rows of wet glass and terracotta pots forming the middle ground. Make one slow, level dolly forward as the leaves part naturally and reveal a gardener opening the far door. Condensation catches warm sunrise light while the floor remains cool and reflective. Maintain straight greenhouse frames and believable depth. No pan, no orbit, no cut. End with the gardener centered beneath the open doorway.
Verify that foreground, aisle, and doorway keep a coherent distance; the dolly should reveal the space without warping frames or jumping the gardener forward.
Coordinate visible motion with a few explicit sound beats while keeping the scene simple enough to audit.
Close three-quarter view of a small espresso machine on a quiet café counter before opening. A barista locks the portafilter with one firm click, presses the brew switch, and two thin streams of coffee begin together. Time the sounds clearly: metal click at the first action, low pump hum immediately after the switch, then a soft cup resonance as coffee lands. Keep the camera fixed and the machine, cup, hands, reflections, and counter geometry stable. No speech, no music, no cut.
Check whether each sound follows its visible cause and whether the streams begin together; the audio toggle is a request control, not a quality guarantee.
Image to Video
4 starter recipes
Reference inputAnimate only the hand, pour, steam, and changing highlights while protecting the still life's carefully placed geometry.
Reference input — prior Wan 2.7 frame; not a Wan 3.0 output.
Use the uploaded image as the first frame and preserve the exact cup, teapot, tray, railing, landscape, colors, and composition. Continue the tea pour with a small natural wrist movement. Let a thin ribbon of steam curl upward and drift slightly toward the valley while sunlight glints gently across the ceramic glaze. Keep the cup and pot shapes rigid, the horizon locked, and the camera almost still with only a very subtle forward ease. No new objects, no cut, no large parallax, no shape change.
Protect the cup rim, handle, teapot spout, railing, and horizon first; motion should remain local to the pour, steam, hand, and light.
Download this reference, then upload it as the first frame. Try prompt switches the mode and fills the text; it does not attach the file.
Reference inputMove the vehicle away from the viewer while treating the road, horizon, and desert masses as stable anchors.
Reference input — prior Wan 2.7 frame; not a Wan 3.0 output.
Use the uploaded image as the first frame. Keep the same car, paint, road markings, mountains, desert palette, and wide composition. The car accelerates gradually away along the road while a low dust trail expands behind the rear wheels, thins in the crosswind, and catches warm side light. Hold the horizon and road perspective steady with a restrained telephoto follow, not a dramatic chase. Preserve vehicle proportions and wheel placement. No extra traffic, no camera roll, no terrain change, no cut.
Check the car scale against the road and confirm that the dust responds to its path without sliding the horizon or reshaping the vehicle.
Download this reference, then upload it as the first frame. Try prompt switches the mode and fills the text; it does not attach the file.
Reference inputAdd energetic sparks and a gentle camera push while locking the worker's equipment and the building's straight lines.
Reference input — prior Wan 2.7 frame; not a Wan 3.0 output.
Use the uploaded image as the first frame and preserve the worker's helmet, gloves, posture, welding tool, steel frame, tables, and workshop layout. The welding arc pulses as bright sparks scatter downward, bounce briefly, and fade before reaching the floor. Add a slow, level camera push of only a few centimeters. Smoke rises in a narrow plume and overhead light flickers subtly on nearby metal. Keep every beam straight and the protective equipment unchanged. No face reveal, no object duplication, no cut, no camera shake.
Inspect helmet and hand continuity, tool contact, spark direction, and beam geometry; reject motion that makes the workshop breathe or bend.
Download this reference, then upload it as the first frame. Try prompt switches the mode and fills the text; it does not attach the file.
Reference inputSeparate foreground rain, distant traffic, and a fixed interior so atmospheric motion does not dissolve the composition.
Reference input — prior Wan 2.7 frame; not a Wan 3.0 output.
Use the uploaded image as the first frame. Keep the window frame, cup, sill, interior reflections, skyline, and overall color balance fixed. New rain beads gather on the glass, merge, and slide downward at different speeds while distant car lights drift softly along the wet street. Let focus breathe once from the nearest droplet toward the city and return. The camera remains locked. Preserve the cup shape and window lines. No lightning, no person entering, no large background movement, no cut.
Check that droplets move on the glass plane, traffic stays distant, and neither the cup nor the window frame drifts during the focus change.
Download this reference, then upload it as the first frame. Try prompt switches the mode and fills the text; it does not attach the file.
Mixed-reference planning
Give each source one reviewable role and state which source wins when references disagree.
| Source | Useful job | Current limit | Conflict check |
|---|---|---|---|
| Images | Protect appearance, identity, geometry, palette, or composition. | Up to 10 files | Name a priority when identity, viewpoint, light, or details differ. |
| Videos | Supply movement, camera behavior, timing, or interaction. | Up to 5 files | Choose compatible direction and tempo; say what to borrow. |
| Audio | Guide rhythm, atmosphere, event timing, or sound density. | Up to 5 files | Avoid several sources competing for the same beat. |
| Prompt | Name each file's role, the action, constraints, and final state. | One written direction | Resolve conflicts explicitly: appearance A, motion B, rhythm C. |
Output planning
Lock delivery settings before polishing the prompt. Runtime and resolution drive the client estimate.
Choose any whole-second duration from 2 through 30. Give short clips one action; divide longer clips into a beginning, change, and ending. Only output duration enters the estimate.
Use adaptive when input images define composition. Otherwise choose 16:9, 9:16, 1:1, 4:3, or 3:4, then write framing details for that destination.
Rates are 15 / 25 / 50 credits per output second. Five seconds estimates 75 / 125 / 250 credits; ten seconds estimates 150 / 250 / 500.
Enable audio only when sound belongs to the brief. The displayed total is a client estimate; the signed-in service response and actual debit are authoritative.
First-task workflow
A controlled first attempt tests one input structure and one set of settings. Change one cause at a time.
Define one shot, its destination, main action, and final state. Remove secondary events that cannot be reviewed inside the selected duration.
Choose text, frames, or mixed references. Upload authorized files, remove duplicates, and state what each surviving input controls.
Select frame, resolution, 2–30 second runtime, and audio. Read the estimate, then hold those choices steady while comparing prompts or references.
Sign in, then record acceptance, status, returned media, debit, and download. That complete live path remains unverified, so inspect observed behavior.
Task fit
Match the input mode to the fact that needs protection. These are preparation patterns, not performance claims.
Start with a clean 9:16 frame when identity or layout already works. Ask for one local movement; review continuity, edge crops, and negative space.
Use a first frame or consistent images to anchor shape, labels, materials, and color. Review silhouette, text, contact points, and reflections first.
Use start and optional end frames when both ends matter. Describe the connecting action; review spatial logic, identity, camera path, and the final transition.
Use audio when rhythm, atmosphere, or a cue is central. Review sound-to-motion timing, or compare another contract in the AI video generator workspace.
Evidence boundary
ZMS has implemented text, frame, and mixed-reference controls, upload handling for supported media, request and status mapping, output settings, audio intent, and a client estimate. These facts show what the page can prepare, not returned-video quality.
Still missing is a complete signed-in run in which the service accepts a task, reaches a terminal status, returns playable media, applies the debit, and permits download. No production readiness, speed, stability, or output-quality claim follows before that evidence exists.
For a batch, save the mode, prompt, inputs, settings, estimate, task ID, result, and balance change. The backend is the final authority; review current terms at ZMS AI pricing.
FAQ
Quick answers about modes, limits, settings, credits, recipe inputs, and verification.
It prepares Wan 3.0 tasks from text, a required first frame with an optional last frame, or mixed references. The controls and adapter exist; complete signed-in delivery and debit remain unverified.
Use text for an open composition, image mode when the first frame protects appearance or layout, and mixed references only when separate files have distinct jobs.
Yes. Image mode requires one first frame and accepts one optional last frame. Add the last frame only for a meaningful ending and describe a plausible transition.
Mixed-reference upload limits are ten images, five videos, and five audio files. At least one file is required, and every file should have a named role.
The interface lists 480p / 720p / 1080p; adaptive, 16:9, 9:16, 1:1, 4:3, and 3:4; 2–30 whole seconds; and generated audio on or off.
The client uses 15 / 25 / 50 credits per output second at 480p / 720p / 1080p. Reference duration is excluded; the backend debit is authoritative.
No. They are labeled prior Wan 2.7 frames offered only as inputs, not Wan 3.0 results. A complete signed-in run through acceptance, result, debit, and download is also still unverified.