13 ratio pairs
Text-to-image offers 13 ratios, each at 1K and 2K.
Plan an Imagine Image 2.0 task from text or one to three reference images, with visible resolution, quality, output-count, format, and credit controls.
Loading generator…
Model identity

On ZMS AI, Grok Image 2.0 is a search phrase for Imagine Image 2.0, not a family page. Its workspace identifier is grok-imagine-image-2.0.
Prepare text-to-image or image-to-image with one to three references. The contract and estimate are visible, but a paid signed-in test must prove submission, returned image URLs, and final debit. This page presents planning material, not results.
Input paths
Start with your material. The ZMS AI image workspace covers broader workflows; this page holds Imagine Image 2.0 task choices.

| Path | Prepare | Confirmed boundary |
|---|---|---|
| Text-to-image | A written brief and one ratio pair. | 26 dimensions: 13 ratios at 1K or 2K. |
| Image-to-image | One to three references and a change request. | Reference editing is 1K/2K without dimension selection. |
| Shared selection rules | Review quality, count, format, and estimate. | Both paths allow low/medium quality and 1–20 outputs. |
Confirmed settings
These cards document the selector boundary; they do not imply hidden controls.
Text-to-image offers 13 ratios, each at 1K and 2K.
With references, choose 1K or 2K; image-to-image removes ratio selection.
Quality is low or medium; a production task still needs review.
Choose 1–20 outputs. The estimate is not final debit.
Text dimensions
Choose one row and one resolution for a text-only task. Reference-led edits show only their 1K or 2K resolution selector instead of these width and height pairs.
| Aspect ratio | 1K dimensions | 2K dimensions |
|---|---|---|
| 1:1 | 1024 × 1024 | 2048 × 2048 |
| 3:4 | 864 × 1152 | 1776 × 2368 |
| 4:3 | 1152 × 864 | 2368 × 1776 |
| 9:16 | 720 × 1280 | 1584 × 2816 |
| 16:9 | 1280 × 720 | 2816 × 1584 |
| 2:3 | 832 × 1248 | 1664 × 2496 |
| 3:2 | 1248 × 832 | 2496 × 1664 |
| 9:19.5 | 576 × 1248 | 1344 × 2912 |
| 19.5:9 | 1248 × 576 | 2912 × 1344 |
| 9:20 | 576 × 1280 | 1440 × 3200 |
| 20:9 | 1280 × 576 | 3200 × 1440 |
| 1:2 | 704 × 1408 | 1456 × 2912 |
| 2:1 | 1408 × 704 | 2912 × 1456 |
Credit estimate
Each row is a per-output estimate. Multiply it by 1–20 outputs, then use ZMS AI credit packs when ready to fund a task; final debit needs E2E verification.
| Resolution | Quality | Credits per output |
|---|---|---|
| 1K | Low | 12 credits |
| 1K | Medium | 18 credits |
| 2K | Low | 18 credits |
| 2K | Medium | 24 credits |
Prompt starters
These are editable planning inputs, not images or samples. For writing guidance, read the GPT Image 2 prompt guide, then select settings here.
Four text-only starting briefs
4 starter recipesDefine object, set, and material.
Studio still life: one brushed steel travel mug on a limestone plinth, three-quarter view, daylight from upper left, grey paper backdrop. Keep one handle, clean rim, realistic highlights, small shadow; no label, extra objects, hands, liquid, or glare.
Check handle count, rim, reflections, shadow, and type.
Reserve a text-safe field without lettering.
Vertical poster background: a cobalt paper wave rises from the lower third on bone white. Keep the upper half open for typesetting, with one shadow, crisp edges, subtle grain, and no words, logos, symbols, people, or extra objects.
Check upper-field space and absent lettering.
Specify light, posture, wardrobe, and frame.
Editorial half-length portrait of an adult ceramic artist, daylight studio, looking just past camera, charcoal shirt, neutral apron, hands at waist. Soft right window light, clay palette, 85mm perspective; no text, logos, extra hands, duplicated tools, distorted shelves, or retouching.
Check hands, shelf geometry, eye direction, and light.
Anchor location, time, and path.
Wide evening scene of one cyclist in a coastal town, viewed from behind. A narrow road curves to low white buildings, sea beyond the roofline, coral sky fading blue. Keep one bicycle and natural perspective; no crowd, signage, text, extra bicycles, impossible architecture, or flare.
Check bicycle count, road, horizon, scale, and sky.
Reference-led editing
When an image-to-image task needs one to three references, write down what each source may influence before you upload it. The useful boundary is not a hidden editing tool; it is a brief that makes the result easier to review after an actual task returns, with clear source roles and review criteria.

Name each source as composition, material cue, or subject constraint; keep every reference role distinct.
List identity, camera position, object count, geometry, or light direction that must stay clearly recognizable.
Describe swap, cleanup, mood, or material adjustment without inventing masks, brushes, search, or seed controls.
List returned-image details carefully to compare so a plausible result is never accepted without inspection.
Suggested edit brief
Use the first uploaded image only as the composition and camera reference: retain the seated adult subject, three-quarter crop, tabletop position, window direction, and visible hand count. Use the second uploaded image only as a material reference for the dark green glazed ceramic cup; keep its subtle speckle and satin reflection, but do not copy logos or lettering. Change the room from a bright daytime café to a quiet early-evening study with a deep blue wall and one warm practical lamp behind the subject. Preserve realistic table perspective, natural fingers, a single cup, and a readable boundary between lamp glow and window shadow. Do not add people, signs, text, duplicate cups, altered facial features, extra hands, dramatic haze, or a different camera angle.
Verification boundary
Only a signed-in paid E2E run can connect the request, returned media, and debit.
| E2E check | What the test must show | What this page does instead |
|---|---|---|
| Task creation | Signed-in submission accepts the selected request and visible settings. | Keep Generate and validation without delivery claims. |
| Status and image URLs | Polling returns task state and image URLs in the library. | Publish no results, speed, quality, or reliability claims. |
| Final debit | Debit matches selected resolution, quality, and output count. | Show an estimate and retain noindex, follow. |
FAQ
These answers separate the search phrase, the specific model workspace, the current request boundary, and the E2E evidence that is still pending.
No. Grok Image 2.0 is the search phrase served by this page. The specific workspace name is Imagine Image 2.0, and the current model identifier is grok-imagine-image-2.0.
No. This page is limited to the Imagine Image 2.0 version workspace. A future /grok-image route can cover family-level choices without duplicating this model-specific contract.
Text-to-image exposes 13 aspect-ratio rows, each with one 1K and one 2K width-by-height pair. That makes 26 confirmed text dimension choices.
Use one to three reference images. In that reference-led path, the workspace exposes 1K or 2K resolution rather than a separate width and height selector.
The current workspace exposes low or medium quality and one through 20 requested outputs. The estimate multiplies the selected per-output rate; final debit still needs E2E evidence.
A signed-in paid task must still prove creation, returned image URLs, and final debit together. Until then, this page does not present a Showcase, performance conclusion, reliability claim, or verified delivery statement.