T3lescope 🔭

Arbitrary-Resolution High-Fidelity Generative
Surface Reconstruction from Images

Atsuhiro Noguchi1, Tianhan Xu1, Yiming Liang1, Yuta Kikuchi1, Masahiro Ishiyama1, Shintaro Takagi2, Hitoshi Murai2, Eiichi Matsumoto1

1Preferred Networks, Inc.  2The University of Tokyo

T3lescope reconstructs the garden scene with a three-level cascade

T3lescope 🔭 is a generative reconstruction method that recovers high-fidelity 3D scene meshes from posed multi-view images without per-scene optimization. It directly outputs a high-resolution mesh with mesh flow models. We adopt a coarse-to-fine cascade: the entire scene is generated at a coarse resolution first, and regions of interest are then upscaled. The cascade layout and the mesh resolution are chosen at inference time. Colors mark the cascade level each surface was generated at. T3lescope 🔭 can reconstruct meshes competitive with those of methods that use per-scene optimization.

Video (1080p resolution recommended)

Abstract

We reconstruct high-fidelity 3D scene meshes from posed multi-view images without per-scene optimization, across scales ranging from single objects to large outdoor scenes. Per-scene optimization methods lack the learned 3D prior needed when observations are sparse or surfaces are glossy or transparent. Existing generative methods leverage such priors to complete geometry in sparsely observed regions, but typically operate at a fixed resolution over a limited spatial extent, trading spatial coverage against detail. Reconstructing a large scene therefore often requires partitioning it into independently processed overlapping local regions, making it difficult to maintain global geometric consistency. To address these issues, we propose T3lescope, which applies a single fixed-resolution generator across scene scales in an inference-time coarse-to-fine cascade. A coarse level establishes the scene layout, and finer levels perturb and denoise geometry inherited from the coarser level within progressively finer spatial cells to recover surface detail. The model is trained on individual cells at multiple scales and shares its weights across all levels, so no hierarchy is fixed during training, and the number of levels, cell scales, and cell locations are determined at inference time. On indoor, outdoor, and city-scale scenes, T3lescope outperforms feed-forward and generative baselines, matches or surpasses per-scene optimization, and recovers fine structures as well as glossy and transparent surfaces. These results show that our method generalizes across diverse scenes, view counts, and image resolutions.

Method

Method overview
Method overview. Ln denotes level n. (a) DA3 points place the L0 cells and select their views. Each finer level halves the cell size and refines the mesh of the level above. (b) Each image is encoded once as a tile pyramid, and from each view a cell takes the tiles at the pyramid level that matches its voxel spacing. (c) All cells at all levels share the SS and mesh generators, which attend to these tiles. At finer levels, the SS stage starts from the noised parent (SDEdit).

Results

Each video compares a mesh reconstructed by a baseline (left) with ours (right), both rendered from the same camera path with the same shading. Drag on a video or image to move the divider; click the round button to play or pause, and the bar below a video to seek.

Mip-NeRF 360 · garden

Input: 185 views. Baseline: GaussianWrapping.

Loading…

In-house capture · five people

Input: 56 views. Baseline: GaussianWrapping.

Loading…

Tanks and Temples · ignatius

Input: 256 views. Baseline: GaussianWrapping.

Our method reconstructs the surfaces within the convex hull of the camera trajectory.

Loading…

Tanks and Temples · caterpillar

Input: 256 views. Baseline: GaussianWrapping.

Our method reconstructs the surfaces within the convex hull of the camera trajectory.

Loading…

ScanNet++ · 578511c8a9

Input: 253 views. Baseline: GenRecon.

Loading…

Single-frame comparison with GenRecon and the real image, at a camera held out from the inputs. Our mesh is more faithful to the image than GenRecon's.

Left side:

City blocks

Input: 256 views. Baseline: CityGaussianV2.

Loading…

BibTeX

@article{noguchi2026t3lescope,
  title   = {{T3lescope}: Arbitrary-Resolution High-Fidelity Generative Surface Reconstruction from Images},
  author  = {Noguchi, Atsuhiro and Xu, Tianhan and Liang, Yiming and Kikuchi, Yuta and Ishiyama, Masahiro and Takagi, Shintaro and Murai, Hitoshi and Matsumoto, Eiichi},
  journal = {arXiv preprint arXiv:2610.03308},
  year    = {2026}
}