Marble 2 beta

Depth-guided generation

atlasChisel generates posed RGB views from optional depth-only context, a required text prompt, a set of target cameras, and optional posed photographs of the scene that the generated views stay consistent with.

Account availability: call POST /api/v2/tasks:atlasChisel only when it appears under API reference for the selected account.

When to use it

Use this task when you have a graybox, blockout, or another geometric layout expressed as posed depth images and want Marble to render coherent RGB views at the same cameras. You can also leave contextFrames empty to generate from the prompt alone.

The request combines:

  • contextFrames: zero or one depth-only frame per target camera, in target order.
  • sourceFrames: zero or more posed RGB photographs of the scene, each an image and its camera.
  • targetCameras: 1 to 32 pinhole cameras for the views to generate.
  • prompt: a required description of the environment, materials, lighting, and style.
  • modelParameters: optional seed and numSteps overrides.
  • voxelSize and voxelGridOrientation: optional controls for quantizing the depth conditioning.

atlasChisel serves one public depth-only contract, so it does not accept a model selector. Start with the documented defaults and change one sampling or voxel control at a time. Keep the prompt, seed, and camera set with the operation so a result can be reproduced.

Camera and depth requirements

Every context and target camera must be a 1280 x 720 pinhole camera. Each context depth raster must also be exactly 1280 x 720 and aligned with its camera. The public endpoint does not resize or crop depth maps for you.

When context is present, provide exactly one frame per target camera in the same order. Each frame carries only:

  • camera
  • depth.depthAsset for a linear-depth EXR, or depth.logDepthAsset for an inverted log-depth PNG

Do not send imageAsset, maskAsset, or depth.confidenceAsset to atlasChisel; the endpoint rejects them instead of silently ignoring them. Keep every camera and depth image in one coordinate convention and scale.

Three.js EXRLoader: If client code decodes or rewrites a linear-depth EXR before upload, account for EXRLoader's bottom-up CPU row order. The task expects depth aligned to the camera's top-left image grid. See Decode linear-depth EXR rows for a reference accessor.

Keep generated views consistent with photographs

When you already have posed photographs of the scene, pass them in sourceFrames. Each entry carries an imageAsset and its camera, the same posed RGB frame shape atlasMasked takes, with no maskAsset and no depth. The model treats them as views it has already committed, so a target camera that sees the same surfaces keeps their materials, colors, and lighting instead of inventing new ones from the prompt.

  • Cameras must be in the same coordinate convention and scale as the target cameras and any context depth, but they need not match a target camera.
  • Any resolution works. Each photograph is cover-resized and center-cropped onto 1280 x 720 before it reaches the model, so keep the subject away from the edges of a photograph with a different aspect ratio.
  • contextFrames, sourceFrames, and targetCameras together fill at most 64 views of the model's rollout. With one depth frame per target, every source photograph costs half a target: 4 photographs leave room for 30 targets.
json
{  "contextFrames": [{ "camera": { "...": "..." }, "depth": { "...": "..." } }],  "sourceFrames": [    {      "imageAsset": { "assetId": "asset_photo_01" },      "camera": {        "extrinsics": {          "position": [1.2, 1.6, -0.4],          "quaternion": [0, 0.38, 0, 0.92],          "coordinateSystem": "rub"        },        "intrinsics": {          "width": 1920,          "height": 1080,          "fx": 1400,          "fy": 1400,          "cx": 960,          "cy": 540        }      }    }  ],  "targetCameras": [{ "...": "..." }],  "prompt": "A stone courtyard with climbing ivy"}

Prompt and voxel controls

prompt must be non-empty whether or not context depth is present. Leave enhancePrompt at true for ordinary prose, or set it to false when the prompt is already in the structured scene-caption form you want the model to receive unchanged.

voxelSize is measured in the normalized scene gauge, where 90th-percentile disparity is approximately 1. Omit it or set it to 0 to keep the original depth conditioning. The recommended range is 0 to 0.4; values around 0.05 to 0.15 are useful starting points for blockouts.

With a nonzero voxelSize, voxelGridOrientation: "world" (the default) aligns the grid to the request's world axes so world-aligned faces quantize flat. "gauge" retains the legacy grid aligned to the recentered mean camera pose.

Handle the result

The task returns a long-running operation. On success, frames contains one generated image and camera per target camera, in request order.

Set returnDepth: true only with at least two target cameras whose centers are not all identical. Returned depth is reconstructed from the generated targets, not copied from the context. A source photograph joins that reconstruction as an anchor, and gets no depth back, only when its image is 1280 x 720 and its camera's principal point is centered. A pure-rotation rig with no anchor can still report success with a depth scale of 0.0 or NaN because it has no translation baseline.

Example completed output

After the operation reports done: true, its task-specific result is in operation.response:

json
{  "frames": [    {      "imageAsset": {        "assetId": "asset_chisel_view_0",        "url": "https://example.com/chisel-view-0.png"      },      "camera": {        "extrinsics": {          "position": [0, 1.6, 0],          "quaternion": [0, 0, 0, 1],          "coordinateSystem": "rub"        },        "intrinsics": {          "width": 1280,          "height": 720,          "fx": 900,          "fy": 900,          "cx": 640,          "cy": 360        }      }    }  ],  "promptUsed": "A warm stone courtyard with climbing ivy",  "requestId": "request_example"}

When returnDepth is true, each generated frame also carries a depth buffer.

Use the Operations polling pattern, and open the atlasChisel reference for the live request schema, defaults, and response types.