Marble 2 beta
Depth-guided generation
atlasChisel generates posed RGB views from optional depth-only context, a
required text prompt, a set of target cameras, and optional posed photographs
of the scene that the generated views stay consistent with.
Account availability: call
POST /api/v2/tasks:atlasChiselonly when it appears under API reference for the selected account.
When to use it
Use this task when you have a graybox, blockout, or another geometric layout
expressed as posed depth images and want Marble to render coherent RGB views at
the same cameras. You can also leave contextFrames empty to generate from the
prompt alone.
The request combines:
contextFrames: zero or one depth-only frame per target camera, in target order.sourceFrames: zero or more posed RGB photographs of the scene, each an image and its camera.targetCameras: 1 to 32 pinhole cameras for the views to generate.prompt: a required description of the environment, materials, lighting, and style.modelParameters: optionalseedandnumStepsoverrides.voxelSizeandvoxelGridOrientation: optional controls for quantizing the depth conditioning.
atlasChisel serves one public depth-only contract, so it does not accept a
model selector. Start with the documented defaults and change one sampling or
voxel control at a time. Keep the prompt, seed, and camera set with the
operation so a result can be reproduced.
Camera and depth requirements
Every context and target camera must be a 1280 x 720 pinhole camera. Each context depth raster must also be exactly 1280 x 720 and aligned with its camera. The public endpoint does not resize or crop depth maps for you.
When context is present, provide exactly one frame per target camera in the same order. Each frame carries only:
cameradepth.depthAssetfor a linear-depth EXR, ordepth.logDepthAssetfor an inverted log-depth PNG
Do not send imageAsset, maskAsset, or depth.confidenceAsset to
atlasChisel; the endpoint rejects them instead of silently ignoring them.
Keep every camera and depth image in one coordinate convention and scale.
Three.js EXRLoader: If client code decodes or rewrites a linear-depth EXR before upload, account for
EXRLoader's bottom-up CPU row order. The task expects depth aligned to the camera's top-left image grid. See Decode linear-depth EXR rows for a reference accessor.
Keep generated views consistent with photographs
When you already have posed photographs of the scene, pass them in
sourceFrames. Each entry carries an imageAsset and its camera, the same
posed RGB frame shape atlasMasked takes, with no maskAsset and no depth.
The model treats them as views it has already committed, so a target camera
that sees the same surfaces keeps their materials, colors, and lighting instead
of inventing new ones from the prompt.
- Cameras must be in the same coordinate convention and scale as the target cameras and any context depth, but they need not match a target camera.
- Any resolution works. Each photograph is cover-resized and center-cropped onto 1280 x 720 before it reaches the model, so keep the subject away from the edges of a photograph with a different aspect ratio.
contextFrames,sourceFrames, andtargetCamerastogether fill at most 64 views of the model's rollout. With one depth frame per target, every source photograph costs half a target: 4 photographs leave room for 30 targets.
Prompt and voxel controls
prompt must be non-empty whether or not context depth is present. Leave
enhancePrompt at true for ordinary prose, or set it to false when the
prompt is already in the structured scene-caption form you want the model to
receive unchanged.
voxelSize is measured in the normalized scene gauge, where 90th-percentile
disparity is approximately 1. Omit it or set it to 0 to keep the original
depth conditioning. The recommended range is 0 to 0.4; values around 0.05 to
0.15 are useful starting points for blockouts.
With a nonzero voxelSize, voxelGridOrientation: "world" (the default)
aligns the grid to the request's world axes so world-aligned faces quantize
flat. "gauge" retains the legacy grid aligned to the recentered mean camera
pose.
Handle the result
The task returns a long-running operation. On success, frames contains one
generated image and camera per target camera, in request order.
Set returnDepth: true only with at least two target cameras whose centers are
not all identical. Returned depth is reconstructed from the generated targets,
not copied from the context. A source photograph joins that reconstruction as
an anchor, and gets no depth back, only when its image is 1280 x 720 and its
camera's principal point is centered. A pure-rotation rig with no anchor can
still report success with a depth scale of 0.0 or NaN because it has no
translation baseline.
Example completed output
After the operation reports done: true, its task-specific result is in
operation.response:
When returnDepth is true, each generated frame also carries a depth
buffer.
Use the Operations polling pattern, and open the atlasChisel reference for the live request schema, defaults, and response types.