On this page

Video workflow

How to Use Veo 3.1: A Practical Shot Workflow

Choose the right Veo 3.1 mode, write one reviewable shot, and keep version-specific controls separate from unverified delivery claims.

Mountain lake frame from an archived Veo 3.1 video example

Archived Veo 3.1 example from a prior internal project. It is shown as a workflow reference, not as a current ZMS AI live-generation test; the source project has no separate rights-clearance ledger.

Choose the Veo 3.1 mode before writing the prompt

Veo 3.1 in ZMS AI is not one uniform set of controls. Lite opens by default, Fast and Pro add options, and Reference mode changes which fields are submitted. Choose the path from the input you actually have instead of drafting a prompt around controls that may disappear after a mode switch.

Use ZMS AI to keep the selected version, input mode, prompt, settings, and service status beside the task. A useful starting decision is simple: text for an open scene, image for a defined opening frame, an optional end frame for a planned arrival, or Reference mode when several assets need to guide identity or place.

The current client adapter is implemented, but credit rates and a signed-in live create-and-check cycle have not been verified. Treat the workspace as a task-preparation surface until a real returned video confirms that the selected provider path accepts the request.

  • Lite: text or image, optional end frame, no audio or Reference mode.
  • Fast or Pro: text or image, optional end frame, audio control, and up to 4K in the UI.
  • Reference: 1–3 images, forced Pro, with a narrower submitted parameter set.
  • Every path: verify account, task status, returned media, and current credit behavior.

Define one shot and its delivery job

Before choosing cinematic language, state what the clip must communicate. Name the audience, destination, aspect ratio, duration, subject, visible action, and final beat. A four- to eight-second task needs one readable event, not a compressed sequence of unrelated scenes.

Write the stable facts separately from the timed event. Identity, product shape, wardrobe, location, weather, and lighting direction may need to remain constant. Action, camera movement, focus change, and environmental motion explain what develops from the opening frame to the ending.

Open the Veo model family when the version itself is undecided. Veo 3 and Veo 3.1 should not be compared by changing model, prompt, ratio, duration, and reference material at the same time. Keep the overlap stable so the result answers a specific production question.

Use Lite as a focused default, not a universal mode

Lite is the first Veo 3.1 version shown in the current ZMS workspace. It prepares text-to-video or image-to-video requests at 720p or 1080p, with 16:9 or 9:16 framing and a duration of 4, 6, or 8 seconds. Image mode can include one start frame and an optional end frame.

The Lite interface hides the audio control and does not offer multi-reference mode. Do not write dialogue or sound effects as if the current Lite request will submit an audio setting. Keep the brief centered on visible action, camera behavior, composition, and continuity.

Lite is useful for a compact first benchmark because it removes several variables. That does not establish a quality or speed ranking. The labels in the client are provider paths, while live delivery, processing time, and cost remain evidence that must come from a completed signed-in task.

Plan Fast and Pro with their additional controls

Fast and Pro retain text mode, one start frame, and an optional end frame. The current interface adds 4K to 720p and 1080p, keeps 4-, 6-, and 8-second choices, and shows an audio toggle in text or image mode. Both also provide a negative-prompt field.

Choose a higher resolution only after the shot direction is stable. More pixels do not fix an unclear event, an unmotivated camera move, or a composition designed for the wrong destination. Run the smallest useful benchmark, then preserve the prompt and settings when moving to another resolution.

Do not promise that Fast is faster or Pro is better until real ZMS tasks provide comparable timing, cost, and output evidence. The adapter maps these labels to different request values, but registration in code is not a service-level benchmark.

Give the start and end frames different jobs

A start frame anchors the opening composition, subject, environment, and visual direction. An end frame defines where the shot should arrive. When both are used, the prompt should explain the action and camera path that connect them instead of restating everything already visible.

Choose compatible frames. A radical change in camera height, subject scale, lighting, or location asks the transition to solve several edits at once. Keep the visual relationship clear, then describe the temporal change: who moves, what the camera does, which environmental details respond, and how the motion settles.

In the Veo 3.1 workspace, the end-frame field appears after a start image is present. Confirm the active version before submission because Lite, Fast, and Pro have different output and audio controls even when the same two frames are visible.

Use Reference mode for 1–3 assets with named roles

Reference mode accepts between one and three images in the current interface and forces the Pro version. Give every image one purpose: character identity, product form, location, prop, or another stable asset. Conflicting images that answer the same question differently create an unclear brief before generation begins.

The current adapter submits the prompt, reference images, resolution, and optional negative prompt for this path. It does not submit aspect ratio, duration, or the audio field. Do not assume that values still visible elsewhere in the generator are silently honored by Reference mode.

Describe the relationship between the assets and the shot. State which subject is primary, what remains recognizable, what action occurs, and how the camera reveals the scene. Avoid asking the model to copy an entire reference when only one object or identity matters.

Write picture, sound, and constraints as separate layers

A practical prompt can follow the viewer’s experience: subject and setting; action over time; camera and composition; light and atmosphere; sound intent when the active mode supports it; then concrete constraints. This order makes each phrase easier to review in the returned clip.

For sound, name the source and timing. Separate dialogue, ambience, music intent, and effects instead of asking for “cinematic audio.” The audio toggle appears only in the current Fast and Pro text or image paths, and a visible switch is not proof that synchronized sound will be returned.

Use the negative prompt sparingly. Protect important facts such as stable product geometry, clean hands, consistent wardrobe, readable screen direction, or the absence of extra people. A long list of unrelated negatives can obscure the positive shot direction.

Review service status before reviewing creative quality

A submitted job must first return a valid task identifier, move through a recognized status, and produce accessible media. Separate connection errors, authentication problems, unsupported parameters, and credit behavior from creative failures in the video itself.

After delivery, watch the full opening, middle, and ending. Check action clarity, identity, anatomy, object geometry, background logic, camera continuity, light, reflections, text, and the resolved final frame. When audio is requested, review dialogue, ambience, synchronization, and unexpected noise separately.

Use the broader AI video workspace when another model or input path better matches the brief. Repeatedly forcing one setup is less informative than a controlled comparison with the same shot objective and acceptance criteria.

  • Connection review: sign-in, task ID, accepted parameters, status, credits, and returned URL.
  • Picture review: action, subject, geometry, camera, continuity, crop, and final frame.
  • Sound review: source, timing, speech clarity, ambience, effects, and synchronization.
  • Production review: rights, consent, destination requirements, compression, and edit handoff.

Save a repeatable recipe and an honest handoff

Keep the approved prompt, input images, selected mode, version, ratio, duration, resolution, audio choice, negative prompt, task identifier, returned file, and review note together. Another editor should be able to understand why the clip passed without reconstructing browser history.

Name the output by project, shot, version, format, and status. Record which imperfections were accepted and which would block publication. Generation approval confirms the creative result; delivery approval separately checks crop safety, compression, playback, sound, and the destination’s technical requirements.

The Google DeepMind Veo model page and Google Cloud Veo 3.1 documentation describe upstream model and endpoint capabilities, while the ZMS page documents this independent interface and adapter. Preserve that distinction in internal notes and public captions. A prior example, a configured control, and a completed live task are three different kinds of evidence.

Continue creating

Turn the guide into a real creative task.

Open Veo 3.1