Marble 2 beta

Images to posed RGBD

images2PosedRGBD estimates camera parameters and linear depth for one or more perspective images. Use it to turn unposed captures into geometry-aware inputs for Marble 2 world workflows.

text
POST /api/v2/tasks:images2PosedRGBD

The task is asynchronous and currently available as a Marble 2 beta endpoint.

Choose your inputs

Supply one or more frames in a stable order. Each imageAsset can be a project-owned asset ID, inline base64 smaller than 10 MiB, or a public HTTPS URL. Assets are the best default for private or reusable media.

json
{  "frames": [    { "imageAsset": { "assetId": "asset_room_01" } },    { "imageAsset": { "assetId": "asset_room_02" } },    { "imageAsset": { "assetId": "asset_room_03" } }  ],  "canonicalizeToFirstFrame": false}

A single image uses the single-view reconstruction path. Multiple images use the multi-view path. The service selects the path from the number of frames.

For multiple images, the service measures focal length on several frames and uses that estimate for guidance. This adds roughly 1-3 seconds to the task and aims to improve camera intrinsics.

Use sharp, overlapping views with stable lighting. Keep the scene mostly static while capturing it, and avoid mixing unrelated places in one request.

Choose the output grid

targetResolution defaults to [1280, 720]. Omit it for the standard Atlas grid, or set another [width, height] when a downstream stage requires a different raster. The returned image, depth, confidence, and camera intrinsics all match that output grid. Pass null to use the reconstruction backend's own default instead.

The task estimates the cameras. Input cameras, masks and depth are not accepted. Gravity alignment and confidence selection are automatic.

Refine the cameras

refinePoses defaults to false. Set it to true to bundle-adjust the estimated cameras against dense matches before answering. The estimated poses are good to about half a degree, which is enough to reconstruct from and not enough to refine against: the residual shows up as blur wherever two views disagree. Turn it on when the output feeds a stage that compares views against each other.

It needs two or more frames, and it costs a matching pass and a solve on top of the reconstruction. A solve that moves the cameras further than the service allows is discarded and the unrefined reconstruction is returned, so a request never fails because the refinement did not hold.

When requested, operation.response.poseRefinementAccepted is true if the refined cameras were kept and false if the original reconstruction was returned. It is null when refinement was not requested.

Get the files in the submission response

base64Outputs defaults to false. A submission can come back already done: true, and by default that response names each output file by assetId and url. Set base64Outputs to true to receive each output file in that response as a base64 data URL instead. The service then skips storing the files before it answers, and you skip downloading them. A later read of the operation names the files by assetId and url either way. See Operations.

Read the result

Persist the operation ID, then use :wait or poll until done is true. The successful response has one frame per input image, in the same order. Each frame includes:

  • the source image, reframed to targetResolution (1280 x 720 by default);
  • a camera describing the returned depth grid;
  • a linear-depth EXR in world units, stored as half-float pixels (relative rounding error up to about 5e-4). Three.js EXRLoader with FloatType and OpenCV decode it to float32; a reader that takes the raw channel gets 16-bit half values.

The camera intrinsics describe the reconstructed grid, which may differ from the original image intrinsics. Unproject the depth with that returned camera rather than with metadata from the source image.

Three.js EXRLoader: Its CPU data array exposes EXR rows bottom-up. Map image-space y to height - 1 - y before unprojecting or matching depth with the returned image and confidence map. See the reference accessor.

Example completed output

After the operation reports done: true, its task-specific result is in operation.response:

json
{  "frames": [    {      "imageAsset": {        "assetId": "asset_room_01",        "url": "https://example.com/room-01.png"      },      "camera": {        "extrinsics": {          "position": [0, 0, 0],          "quaternion": [0, 0, 0, 1],          "coordinateSystem": "rub"        },        "intrinsics": {          "width": 1280,          "height": 720,          "fx": 900,          "fy": 900,          "cx": 640,          "cy": 360        }      },      "depth": {        "depthAsset": {          "assetId": "asset_depth_0",          "url": "https://example.com/depth-0.exr"        },        "confidenceAsset": {          "assetId": "asset_confidence_0",          "url": "https://example.com/confidence-0.png"        }      }    }  ]}

Open the API reference for the exact camera, depth, and operation schemas.