Marble 2 beta

Inpainting

atlasMasked completes masked regions across posed views. Each context frame contains an RGB image, a separate grayscale mask, and its camera. targetCameras declares the views at which completed frames are returned.

Access-dependent preview: only call this task when POST /api/v2/tasks:atlasMasked appears in the API reference for your selected account. It is not part of the default external beta surface yet.

Understand the mask

The mask is an 8-bit grayscale PNG on the same pixel grid as its image:

  • 255 (or normalized 1) keeps the observed pixel;
  • 0 marks the region the model should fill;
  • intermediate values form a soft transition.

Do not apply the mask to the RGB image before uploading it. Send the image and mask as separate assets. Pre-masking removes context the model needs to produce a clean boundary.

Prepare consistent views

Use between one and sixteen well-overlapped views as the practical target for this preview model. Pair each context frame with the target camera at the same index. The API accepts different counts or off-pair cameras, but the partial checkpoint was trained with one context view per target at the same viewpoint, so those requests can produce degraded results.

When removing or replacing an object, make the zero-valued region fully cover that object in every view. Give the mask a small margin. A mask that leaves part of the old object visible often produces remnants or seams.

For a controlled replacement workflow:

  1. Choose a clear anchor frame that faces the object directly and is not too distant or top-down.
  2. Edit that anchor with a strong image editor so it shows the desired result.
  3. Send the edited anchor as contextFrames[0] with an all-ones mask, and its matching camera as targetCameras[0]. Current preview results can depend on this order.
  4. Send the remaining original views with zero-valued replacement regions.

Cluttered backgrounds and weak cross-view coverage make completion harder.

Prepare inputs locally

Use the open-source Marble Multiview Inpainting tools to prepare consistent masks from one anchor edit. Supply the original posed RGBD views, an edited anchor image, and a white-on-black edit-region mask. White in this input mask marks the edited region; the exported API masks use the keep-mask convention described above.

The tools project the anchor region into the other views using depth and cameras, check visibility against target-view depth, and resize/crop images and masks to 1280 × 720 with matching camera intrinsics. Outputs include grayscale keep masks, an all-white mask for the edited anchor, preview overlays, and a request template.

Follow the clone-and-run quickstart to try the included 3D example, then use the input specification for your own scene. Preparation runs locally without an API key or GPU. If you do not have depth, use the repository's manual-mask mode.

Inspect the preview overlays before uploading; projected masks can need manual correction around occlusions or newly exposed surfaces. Upload the prepared RGB images and keep masks using the Assets flow, then replace the asset-ID placeholders in atlas-masked-request.template.json. The tools prepare inputs only; submit the request to atlasMasked to generate completed views.

Match camera and image dimensions

Every context and target camera must declare the exact dimensions the preview model serves: 1280 × 720. An off-size camera is rejected before inference. The server can resample an image and mask onto their camera grid, but it does not perform a geometry-aware crop. Keep the image and mask on the same pixel grid and aspect ratio so the content is not stretched and the masked region stays aligned.

Build the request

json
{  "contextFrames": [    {      "imageAsset": { "assetId": "asset_view_01" },      "maskAsset": { "assetId": "asset_mask_01" },      "camera": {        "extrinsics": {          "position": [0, 1.6, 0],          "quaternion": [0, 0, 0, 1],          "coordinateSystem": "rub"        },        "intrinsics": {          "width": 1280,          "height": 720,          "fx": 900,          "fy": 900,          "cx": 640,          "cy": 360        }      }    }  ],  "targetCameras": [    {      "extrinsics": {        "position": [0, 1.6, 0],        "quaternion": [0, 0, 0, 1],        "coordinateSystem": "rub"      },      "intrinsics": {        "width": 1280,        "height": 720,        "fx": 900,        "fy": 900,        "cx": 640,        "cy": 360      }    }  ],  "prompt": "Add a low wooden bench against the stone wall",  "enhancePrompt": true,  "returnDepth": false}

Write the prompt

With prompt enhancement on (the default), write the prompt as an edit instruction: what to add, remove, replace, or restyle, and where. For example, "replace the yellow car with a stone fountain in the middle of the courtyard". The enhancer reads the instruction together with your original images and writes a detailed description of the finished scene for the model, returned as promptUsed. It does not see the masks, so place the change relative to things in the scene. Concrete appearance details help the result stay consistent across views.

State removals explicitly, such as "remove the parked car". Without a prompt, the enhancer describes the images as they are, masked objects included, and the model tends to put those objects back.

Keep the instruction even when you send an edited anchor frame. Your other views still show the old content, and without the instruction the enhancer describes the set as a change over time, old object included.

Set enhancePrompt: false only to supply that description yourself. The model then receives your text unchanged as a caption of the finished image, so describe the whole finished scene rather than giving an instruction.

Request depth only when useful

Set returnDepth: true to reconstruct completed frames together and attach a depth buffer to each result. It needs at least two target cameras. Pure-rotation camera rigs share a camera center and cannot reliably recover metric scale, so use translated viewpoints when depth scale matters.

Example completed output

After the operation reports done: true, its task-specific result is in operation.response:

json
{  "frames": [    {      "imageAsset": {        "assetId": "asset_completed_view_0",        "url": "https://example.com/completed-view-0.png"      },      "camera": {        "extrinsics": {          "position": [0, 1.6, 0],          "quaternion": [0, 0, 0, 1],          "coordinateSystem": "rub"        },        "intrinsics": {          "width": 1280,          "height": 720,          "fx": 900,          "fy": 900,          "cx": 640,          "cy": 360        }      }    }  ],  "promptUsed": "A low wooden bench against the stone wall",  "requestId": "request_example"}

When returnDepth is true, each completed frame also carries a depth buffer.

Runtime grows with resolution, view count, sampling steps, and optional depth reconstruction. Submit once, retain the operation ID, and use the normal Operations lifecycle rather than resubmitting while work is running.