[论文] Zero-Shot Novel Depth Synthesis Using 3D Foundation Models Scene Repre…

## 论文概要 **研究领域**: CV **作者**: Denis M. Akola, David F. F...

论文概要

研究领域: CV 作者: Denis M. Akola, David F. Fouhey 发布时间: 2026-09-06 arXiv: 2509.04284

中文摘要

3D基础模型(3DFM)如VGGT最近通过前馈Transformer预测丰富的统一表示,推动了3D视觉的边界。这些模型学到的场景表示使其在多个3D视觉任务上表现优异。本文研究如何利用其内部表示从新视角推断场景中的3D信息。我们的假设是:为了解决3D重建任务,这些模型需要学习一种包含大量关于3D场景通用知识的表示。在展示了可以从3DFM内部表示解码隐藏表面后,我们提出了一种名为Z3D的方法,通过对3DFM表示进行潜在扩散来估计未见过视角的点图。我们证明Z3D能够在多个数据集上为新视角预测逼真的深度图。

原文摘要

3D Foundation Models (3DFMs) such as VGGT have recently pushed the boundaries of 3D vision by predicting rich unified representations with feed-foward transformers. The scene representations learned by these models enable strong performance on multiple 3D vision tasks. In this paper, we investigate using their internal representations to infer 3D in the scene from new views. Our hypothesis is that in order to solve the task of 3D reconstruction, these models need to learn a representation that includes a large amount of general knowledge about 3D scenes. After showing that it is possible to decode hidden surfaces from internal 3DFM representations, we propose a method, Z3D, that estimates pointmaps in unseen views by doing latent diffusion on 3DFM representation. We show that Z3D can predi…

自动采集于 2026-09-07

#论文 #arXiv #CV #小凯

发表回复

人生梦想 - 关注前沿的计算机技术 acejoy.com 🐾 步子哥の博客 🐾 背多分论坛 🐾 借一步网 🐾 智柴网 沪ICP备2024052574号-1