SIGGRAPH ASIAACM ToG2026

DecomVoxel

Harnessing 3D-Native Priors with Guided In-situ Denoising Optimization for Decompositional Scene Reconstruction

Junfeng Ni1,2,* Zirui Zhou1,2,* Yixin Chen2,†,✉ Yu Liu1,2 Nan Jiang3 Zhifei Yang3 Song-Chun Zhu1,2,3 Siyuan Huang2,✉

1 Tsinghua University2 State Key Laboratory of General Artificial Intelligence, BIGAI3 Peking University

* Equal contribution† Project lead✉ Corresponding author

Explore the scene

Abstract

Introduction

We propose DecomVoxel, a framework that integrates guided 3D-native priors to enhance decompositional scene reconstruction. By formulating object completion as guided in-situ optimization, our method achieves high-quality topology, geometry, and appearance for both individual objects and backgrounds while strictly preserving original spatial layouts.

Project video Guided in-situ optimization.

Interactive reconstruction / Direct manipulation

Reach into the scene.

Preparing scene Loading 3D reconstruction…
ObjectDrag to interact

Scene explorer / Three rendering modes

One view. Three readings.

Drag both dividers to inspect grid structure, decomposed object color, and final texture.

Method / Guided in-situ reconstruction

Method Overview

Our framework bridges 3D generative priors with neural reconstruction through a guided in-situ denoising optimization. We segment objects from the initial scene and then sequentially optimize their geometry and appearance under adaptive spatial guidance. This process recovers missing structures and consistent textures while preserving the original scene layout, producing high-fidelity topology, geometry, and appearance.

DecomVoxel method overview from initial reconstruction through geometry and appearance optimization to outputs
Pipeline overview Background completion and object-level geometry and appearance refinement.
01 / Preliminaries

Preliminaries

Sparse Voxel Representation

The scene is stored as an octree of sparse voxels carrying position, scale, density, and appearance coefficients. Multi-view instance masks are fused to obtain each object’s voxel subset, while GeoSVR’s geometric uncertainty marks reliable voxels for spatial anchoring.

3D-Native Prior

TRELLIS supplies structure generation, latent appearance refinement, and VAE decoding, allowing incomplete object voxels to become multi-view-consistent geometry and appearance.

02 / Denoise

In-situ Denoising Optimization

We complete each incomplete object directly in the scene coordinate frame by encoding its voxel segment into a latent and iteratively refining it with a pretrained 3D flow-matching prior. An epsilon-based distillation objective adapts update strength along the denoising trajectory, suppressing unreliable high-noise changes while preserving detail as the latent converges.

03 / Anchor

Adaptive Spatial Guidance

To regularize the stochasticity of the 3D-native prior, we retain regions with high reconstruction confidence and inject more generative prior into low-confidence regions, balancing structural preservation with completion.

04 / Pipeline

Optimization Pipeline

01 / Decompositional reconstruction

GeoSVR produces a global sparse voxel scene, instance masks partition objects and background, and geometric uncertainty initializes reliable anchors.

02 / Background completion

Unobserved colors are inpainted and planar depth constraints support re-optimization before extracting the background mesh.

03 / Object refinement

A structural grid recovers complete geometry; DINOv2 features back-projected from multiple views initialize appearance latents for texture synthesis.

Results

Qualitative Comparison

Browse four reconstructed scenes and compare semantic geometry with final appearance.

Analysis

Quantitative comparison

Our method consistently outperforms state-of-the-art baselines across both reconstruction and rendering metrics.

Reconstruction

CD ↓

17.5510.5912.876.5915.493.95
Replica
17.1112.7515.359.7711.754.34
ScanNet++
Rendering

PSNR ↑

18.0418.8918.83–18.5824.21
Replica
16.4016.4617.41–17.5921.71
ScanNet++

Citation

If you find our work useful, please cite DecomVoxel.

BibTeX
@article{ni2026decomvoxel,
  title = {DecomVoxel: Harnessing 3D-Native Priors with Guided In-situ Denoising Optimization for Decompositional Scene Reconstruction},
  author = {Ni, Junfeng and Zhou, Zirui and Chen, Yixin and Liu, Yu and Jiang, Nan and Yang, Zhifei and Zhu, Song-Chun and Huang, Siyuan},
  journal = {ACM Transactions on Graphics},
  year = {2026}
}