Marble 2 beta

Text to image

atlasTextToImage generates image inputs from a text prompt for Marble 2 world workflows.

Access-dependent preview: only call this task when POST /api/v2/tasks:atlasTextToImage appears in the API reference for your selected account. The endpoint is available to enabled Marble 2 beta accounts, and its request may change while the model is being updated.

When to use it

Use atlasTextToImage when you want a convenient world-oriented starting image inside the Marble pipeline. Describe a coherent place, its layout, materials, lighting, time of day, and camera viewpoint.

json
{  "prompt": "A quiet stone courtyard at blue hour, arched walkways, warm window light, wide eye-level view",  "aspectRatio": "16:9",  "seed": 7,  "numSamples": 1}

Optional controls include aspectRatio, enhancePrompt, seed, numSteps, and numSamples. Start with defaults, change one control at a time, and record the prompt, resolved model, and seed with downstream operations.

When numSamples is greater than 1, each image uses independent noise. If seed is provided, sample i uses seed + i. With enhancePrompt set to false, repeating the request reproduces the full set.

aspectRatio accepts 16:9, 9:16, 4:3, 3:4, or 1:1, and falls back to 1:1 when omitted.

Prompt enhancement is enabled by default. Set enhancePrompt to false when you need the submitted wording to reach the model unchanged.

Image quality guidance

Prompt for a world:

  • Describe one navigable environment rather than a collage of unrelated subjects.
  • State spatial relationships: what is foreground, background, left, right, enclosed, or open.
  • Prefer a clear, wide viewpoint with visible surfaces and depth cues.
  • Avoid prominent text, split screens, crowds, fast motion, and large close-up people when the image will feed a world model.
  • Higher detail prompts lead to better generations.

Marble models are strongest on worlds and environments. Humans and dynamic objects are not yet their strongest subjects; see Best practices.

Enterprise customers with a repeated visual domain can work with World Labs to fine-tune and improve our text-to-image model for their specific use case. Contact your World Labs account representative with representative prompts, desired outputs, and evaluation criteria.

For downstream Atlas tasks such as Preview, Chisel, and Reconstruct, higher-quality input images generally produce better results. You can compare images from different providers to evaluate the resulting world quality.

Handle the operation

The task returns a long-running operation. A successful terminal response carries one frame per generated image, the prompt inference actually consumed, and the resolved aspect ratio. Treat URLs according to the returned asset contract and persist the operation ID for traceability.

Example completed output

After the operation reports done: true, its task-specific result is in operation.response:

json
{  "frames": [    {      "imageAsset": {        "assetId": "asset_example",        "url": "https://example.com/generated-0.png"      }    }  ],  "promptUsed": "A quiet stone courtyard at blue hour with coherent architectural depth",  "aspectRatio": "16:9"}

When returnDepth is true, every frame also carries the camera its depth was estimated under and a depth buffer.

Always confirm the live request and response in the API reference before shipping an integration.