Marble 2 beta
Images to posed RGBD
images2PosedRGBD estimates camera parameters and linear depth for one or
more perspective images. Use it to turn unposed captures into geometry-aware
inputs for Marble 2 world workflows.
The task is asynchronous and currently available as a Marble 2 beta endpoint.
Choose your inputs
Supply one or more frames in a stable order. Each imageAsset can be a
project-owned asset ID, inline base64 smaller than 10 MiB, or a public HTTPS
URL. Assets are the best default for private or reusable media.
A single image uses the single-view reconstruction path. Multiple images use the multi-view path. The service selects the path from the number of frames.
For multiple images, the service measures focal length on several frames and uses that estimate for guidance. This adds roughly 1-3 seconds to the task and aims to improve camera intrinsics.
Use sharp, overlapping views with stable lighting. Keep the scene mostly static while capturing it, and avoid mixing unrelated places in one request.
Choose the output grid
targetResolution defaults to [1280, 720]. Omit it for the standard Atlas
grid, or set another [width, height] when a downstream stage requires a
different raster. The returned image, depth, confidence, and camera intrinsics
all match that output grid. Pass null to use the reconstruction backend's
own default instead.
The task estimates the cameras. Input cameras, masks and depth are not accepted. Gravity alignment and confidence selection are automatic.
Refine the cameras
refinePoses defaults to false. Set it to true to bundle-adjust the
estimated cameras against dense matches before answering. The estimated poses
are good to about half a degree, which is enough to reconstruct from and not
enough to refine against: the residual shows up as blur wherever two views
disagree. Turn it on when the output feeds a stage that compares views against
each other.
It needs two or more frames, and it costs a matching pass and a solve on top of the reconstruction. A solve that moves the cameras further than the service allows is discarded and the unrefined reconstruction is returned, so a request never fails because the refinement did not hold.
When requested, operation.response.poseRefinementAccepted is true if the
refined cameras were kept and false if the original reconstruction was
returned. It is null when refinement was not requested.
Get the files in the submission response
base64Outputs defaults to false. A submission can come back already
done: true, and by default that response names each output file by assetId
and url. Set base64Outputs to true to receive each output file in that
response as a base64 data URL instead. The service then skips storing the files
before it answers, and you skip downloading them. A later read of the
operation names the files by assetId and url either way. See
Operations.
Read the result
Persist the operation ID, then use :wait or poll until done is true. The
successful response has one frame per input image, in the same order. Each
frame includes:
- the source image, reframed to
targetResolution(1280 x 720 by default); - a camera describing the returned depth grid;
- a linear-depth EXR in world units, stored as half-float pixels (relative
rounding error up to about 5e-4). Three.js
EXRLoaderwithFloatTypeand OpenCV decode it to float32; a reader that takes the raw channel gets 16-bit half values.
The camera intrinsics describe the reconstructed grid, which may differ from the original image intrinsics. Unproject the depth with that returned camera rather than with metadata from the source image.
Three.js EXRLoader: Its CPU
dataarray exposes EXR rows bottom-up. Map image-spaceytoheight - 1 - ybefore unprojecting or matching depth with the returned image and confidence map. See the reference accessor.
Example completed output
After the operation reports done: true, its task-specific result is in
operation.response:
Open the API reference for the exact camera, depth, and operation schemas.