Marble 2 beta
Text to image
atlasTextToImage generates image inputs from a text prompt for Marble 2 world
workflows.
Access-dependent preview: only call this task when
POST /api/v2/tasks:atlasTextToImageappears in the API reference for your selected account. The endpoint is available to enabled Marble 2 beta accounts, and its request may change while the model is being updated.
When to use it
Use atlasTextToImage when you want a convenient world-oriented starting image
inside the Marble pipeline. Describe a coherent place, its layout, materials,
lighting, time of day, and camera viewpoint.
Optional controls include aspectRatio, enhancePrompt, seed,
numSteps, and numSamples. Start with defaults, change one control at a
time, and record the prompt, resolved model, and seed with downstream
operations.
When numSamples is greater than 1, each image uses independent noise. If
seed is provided, sample i uses seed + i. With enhancePrompt set to
false, repeating the request reproduces the full set.
aspectRatio accepts 16:9, 9:16, 4:3, 3:4, or 1:1, and falls back to
1:1 when omitted.
Prompt enhancement is enabled by default. Set enhancePrompt to false when
you need the submitted wording to reach the model unchanged.
Image quality guidance
Prompt for a world:
- Describe one navigable environment rather than a collage of unrelated subjects.
- State spatial relationships: what is foreground, background, left, right, enclosed, or open.
- Prefer a clear, wide viewpoint with visible surfaces and depth cues.
- Avoid prominent text, split screens, crowds, fast motion, and large close-up people when the image will feed a world model.
- Higher detail prompts lead to better generations.
Marble models are strongest on worlds and environments. Humans and dynamic objects are not yet their strongest subjects; see Best practices.
Enterprise customers with a repeated visual domain can work with World Labs to fine-tune and improve our text-to-image model for their specific use case. Contact your World Labs account representative with representative prompts, desired outputs, and evaluation criteria.
For downstream Atlas tasks such as Preview, Chisel, and Reconstruct, higher-quality input images generally produce better results. You can compare images from different providers to evaluate the resulting world quality.
Handle the operation
The task returns a long-running operation. A successful terminal response carries one frame per generated image, the prompt inference actually consumed, and the resolved aspect ratio. Treat URLs according to the returned asset contract and persist the operation ID for traceability.
Example completed output
After the operation reports done: true, its task-specific result is in
operation.response:
When returnDepth is true, every frame also carries the camera its depth was
estimated under and a depth buffer.
Always confirm the live request and response in the API reference before shipping an integration.