480p or 720p
Choose 480p for direction tests or 720p after motion and framing pass. This route has no 1080p or 4K.
Prepare Seedance 2.5 video tasks from text, image or video references, or a first-and-last-frame path, with 480p or 720p and 4–30 second duration.
Loading workspace…
Model overview
The Seedance 2.5 AI Video Generator in ZMS AI prepares text, first/last-frame, image-reference, or video-reference tasks with one reviewable control set.

ByteDance describes broader joint audio-video generation. ZMS exposes 480p/720p, 4–30 seconds, audio, and listed ratios, but not seed, camera, fast, or think switches.
Compare the Seedance model family first; visible ZMS controls define this route.
Workspace controls
Choose settings from the final destination and keep them stable during prompt revisions.
Choose 480p for direction tests or 720p after motion and framing pass. This route has no 1080p or 4K.
Choose 4–30 seconds in one-second steps and leave one action enough time to resolve.
Use text, a required first and optional last frame, images, or videos. Do not mix image and video references. Text mode lists six ratios plus adaptive; frame mode is adaptive.
The registered driver prepares tasks, but signed-in delivery and debit still require verification.
Multimodal task recipes
These original suggested briefs cover all four input paths and are not recovered outputs. Try preserves mode; supply authorized inputs.
Text, frame anchors, image references, and video references
4 starter recipesBuild three connected beats with one subject and final state.
Create one continuous twenty-second sequence in a bright bookbinding studio. An adult binder folds one cream signature, sews it with blue thread, then places the completed section beneath a small wooden press. Keep the same person, hands, linen apron, paper stack, needle, thread color, press, table, window, and daylight direction through all three beats. Use one slow camera arc from front-left to side view without crossing the work axis. Include only quiet paper movement and room tone; no speech or music. No cut to another room, duplicate tools, extra fingers, changing clothing, warped paper, reversed thread path, logo, text, or late camera jump. End with both hands clear of the closed press.
Check causal beats, identity, tools, paper, camera, sound, and final hands.
Bridge approved opening and closing frames with one plausible transition.
Move continuously from the supplied first frame of an empty museum corridor to the supplied last frame with the same adult visitor standing beneath the far skylight. Preserve corridor width, stone joints, artwork positions, skylight geometry, visitor identity, dark coat, shoe color, and cool daylight. The visitor enters naturally from the right, walks along the center line, and stops at the exact final position while the camera makes one restrained forward track. No cut, morph, teleport, new artwork, changing architecture, camera roll, lighting shift, extra person, or alteration to either anchor composition. End only when the last frame is matched.
Compare anchors, path, identity, artwork, architecture, track, and ending position.
Give each authorized image one subject, product, material, or environment role.
Create one twelve-second restrained product shot from the supplied authorized image references. Reference one controls the compact radio silhouette and front control layout. Reference two controls brushed-aluminum material and grain. Reference three controls the pale limestone studio and window-light direction. Reference four controls only the adult presenter's navy sleeve and hand position. Place one radio on a low plinth while the presenter rotates it a quarter turn and releases it; use a slow product-height camera orbit. Preserve every control, grille slot, proportion, material direction, plinth edge, sleeve, hand count, and room geometry. No copied logo, duplicate product, redesign, extra accessory, mixed identities, cut, text, or new background. End on a clean three-quarter product view.
Audit reference roles, geometry, material, controls, objects, branding, hands, and camera.
Transfer timing or camera behavior while excluding source content and identity.
Use the supplied authorized video references only for two motion roles: reference one contributes the timing of a four-second cloth unfolding action; reference two contributes a slow level camera move from medium to close. Create a different scene with an adult museum preparator unfolding one gray conservation cloth across an empty oak table in a white archive room. Preserve the new subject, gloves, cloth size, table edges, shelf geometry, neutral light, and camera axis. Do not copy any person, clothing, room, object, logo, text, color palette, audio, or final frame from either source. No extra hand, cloth morph, cut, zoom jump, moving shelves, speech, or music. End with the cloth flat and both gloved hands clear.
Confirm only timing and camera transfer; inspect every stated exclusion.
Capability boundary
Separate official scope from ZMS controls without treating 2.5 as a universal upgrade.
| Layer | Official Seedance 2.5 | Current ZMS UI | Review note |
|---|---|---|---|
| Official model scope | Audio-video generation for 30-second stories with references | Broader upstream model scope | Model context, not proof of ZMS controls |
| Text-to-video path | Text-led generation from a structured prompt | ZMS sends structured text to the adapter | Define one event, subject, setting, and camera path |
| Reference-to-video path | References for subjects, scenes, and motion | Up to 30 images or 10 videos | Choose images or videos, not both |
| First-and-last-frame and audio | First/optional last frame plus audio | Required first frame, optional last frame, audio; adaptive framing | No seed/camera/fast/think controls; 2.0 reaches 1080p/4K |

Planning illustration
Use a still to align subject, framing, motion, and review criteria; it is not generated proof.
Bright storyboard planning illustration created for this guide. It is not a Seedance 2.5 output and is not evidence of Seedance 2.5 quality or capabilities.Practical workflow
Use four stages; compare Seedance 2.0 for 1080p/4K or 4–15 seconds.
Name destination, subject, action, and feeling; split unrelated scenes.
Choose text, frame anchors, images, or videos; never mix both reference types.
Lock 480p/720p, 4–30 seconds, and framing before comparison.
Check opening, midpoint, ending, identity, geometry, camera, timing, and resolution.
Production review
Seedance 2.5 and 2.0 expose different tradeoffs, while upstream capability and ZMS delivery remain separate claims. This ledger turns input choice, reference authority, duration, resolution, audio, rights, and service verification into explicit actions before any longer or multimodal task is approved.
| Decision | Current fact | Review action |
|---|---|---|
| Choose the input authority | ZMS exposes text, first and optional last frames, up to 30 images or 10 videos, and audio. Image and video references cannot be mixed in one task. | Assign one job to every authorized source, remove conflicts, preserve exclusions, and compare version controls on the Seedance family before uploading a large set. |
| Choose 2.5 or 2.0 | Seedance 2.5 offers 4–30 seconds, reference paths, first-last frames, audio, 480p/720p, and adaptive framing. Seedance 2.0 keeps 4–15 seconds and reaches 1080p/4K. | Use Seedance 2.0 for its higher resolution path; use 2.5 only when longer timing or its multimodal input authority matters more. |
| Approve production use | The registered driver and visible controls do not prove signed-in completion, returned picture or audio, or final debit. The storyboard is decor, not Seedance 2.5 evidence. | Run one account task, retain inputs and settings, review every frame and sound, confirm consent and rights, then check ZMS AI pricing or the AI video workspace. |
FAQ
Answers about ZMS inputs, framing, duration, model scope, and the planning illustration.
It is a registered ZMS workspace with text, frame anchors, image or video references, audio, 480p/720p, and 4–30 seconds. Delivery remains unverified.
Yes. The first frame is required, the last is optional, and framing is adaptive. Other input paths remain available.
ZMS lists 480p/720p. Text mode offers six ratios plus adaptive; first/last-frame mode is adaptive.
Choose whole seconds from 4–30 and leave one event enough time to resolve.
No. ZMS exposes the listed input paths and audio, but not seed, camera, fast, or think controls.
No. It is decor, not a Seedance 2.5 output or quality evidence.