Marble 2 beta
Generate, pose, and edit a world image
This tutorial chains three Marble 2 tasks into one reproducible workflow:
atlasTextToImage → images2PosedRGBD → atlasMasked
You will generate a world image, estimate its camera, and then edit a masked region while preserving the unmasked scene. The intermediate image and camera are explicit artifacts, which makes the workflow easy to inspect, retry, and automate with a coding agent.
Access-dependent tutorial: all three task endpoints must appear in the API reference for your selected account. At the time of writing,
atlasTextToImageis available to enabled Marble 2 beta accounts andatlasMaskedremains an internal preview. IfatlasMaskedis not enabled for your account, submitting it returns HTTP 404. Complete the first four steps and stop before the masked edit.
Before you start
Create a project API key with tasks.create, operations.read,
assets.create, and assets.read, then configure:
The examples use curl, jq, OpenSSL, and ImageMagick. Install the
dependencies once:
ImageMagick 7 calls its CLI magick; some Linux packages still ship
ImageMagick 6 as convert. Select whichever command is installed:
Keep every JSON response while developing; operation IDs are the audit trail between stages.
Define helpers that retry transient API failures, generate a unique idempotency key for each intentional submission, print progress, and stop polling after the overall timeout. Automatic retries of one submission reuse the same key.
The request helper retries transport failures and transient HTTP responses such
as 429, 500, 502, 503, and 504 with exponential delay capped at eight seconds.
Each response is kept in a temporary file so failed response bodies cannot be
combined with successful JSON from a later attempt. If polling hits the overall
timeout, the operation continues server-side and the printed
wait_for_operation command resumes watching it without submitting another
task. See Operations and Errors and retries.
1. Generate a world image
Ask for a coherent environment with a wide, eye-level view and visible depth cues:
One sample is intentional. Multiple atlasTextToImage samples are variations,
not synchronized views of one scene, so do not treat them as a multi-view
capture.
Read Text to image for prompting and model-quality guidance.
2. Prepare a supported inpainting frame
atlasMasked accepts 1280 × 720 cameras. Resize and center-crop before estimating the
pose so every later artifact shares the same aspect ratio:
Create an 8-bit grayscale mask on the same 1280 × 720 pixel grid. White keeps the original pixels; black marks pixels the model should replace. This sample regenerates a band at the top of the image:
For object removal, paint the object black in every affected view and include a small margin around it.
For multiview edits with existing posed RGBD inputs, use the open-source Marble Multiview Inpainting tools to project an anchor edit into the other views. See Prepare inputs locally for the input format, preview checks, and upload handoff. The single-image example here does not require these tools.
3. Upload the prepared image and mask
Inline asset creation is convenient for PNG files smaller than 10 MiB. For larger inputs, use the upload flow in Assets.
The base64 bytes travel through standard input instead of a shell argument, so multi-megabyte images do not hit the operating system's argument-size limit.
4. Estimate the camera
Run images2PosedRGBD on the prepared image. A single frame selects the
single-image reconstruction path:
estimated-depth.exr is the generated linear-depth map in world units. Keep
the EXR for geometry processing; the image above is a normalized visualization
of this kind of depth data, not the raw floating-point values. The response
also includes per-pixel confidence at
.frames[0].depth.confidenceAsset.url.
If Three.js EXRLoader reads the EXR for CPU geometry processing, apply the
EXR row mapping
before matching depth pixels with the image or confidence map. The downloaded
EXR itself does not need to be rewritten.
The request puts the returned image, depth, confidence, and camera on the same
1280 × 720 crop required by atlasMasked. Copy the camera without changing its
intrinsics:
Do not resize the posed image or depth after this point. A client-side intrinsics scale would miss the center-crop offset already applied by the posing task.
5. Complete the masked region
First confirm that atlasMasked appears under Task endpoints in the live
API reference. If it is absent, the HTTP 404 response is
the expected access-gate behavior, not a problem with your request. Ask your
World Labs contact for preview access or stop here.
If it is available, combine the prepared RGB asset, mask asset, and normalized camera:
This request sets enhancePrompt: false, so the model receives the prompt
exactly as written and reads it as a caption of the finished image, not as
instructions. To keep part of the scene unchanged, leave it white in the mask.
For your own edits, leave enhancement on and write the change as an
instruction, as described in Write the prompt.
The successful response preserves the camera with the completed image. Keep the response, request, and all three operation IDs together so you can trace exactly which generated image and pose produced the edit.
Extend the workflow
- Set
returnDepth: trueonly after adding at least two translated context views; pure-rotation cameras cannot establish reliable metric scale. - Feed multiple real, overlapping captures to
images2PosedRGBDinstead of independent text-to-image variations when you need a multi-view edit. - Continue with
atlasChiselonly after preparing 1280 x 720 depth-only frames and cameras together; posed RGBD output cannot be passed through verbatim. Omit its RGB, mask, and depth confidence assets. - Validate every request against the live API reference, because preview schemas can change during the Marble 2 beta.