论文概要
研究领域: CV 作者: Wenhao Li, Xueying Jiang, Quanhao Qian 发布时间: 2026-07-25 arXiv: 2507.20491
中文摘要
尽管视觉-语言模型(VLMs)进展迅速,但基于2D视觉输入构建的大多数现有模型在处理需要细粒度空间理解和推理的3D任务时往往力不从心。为弥合这一差距,我们提出了VLM-IE3D,这是一个统一框架,通过为VLMs配备从RGB视频中学习的隐式和显式3D几何信息来增强其3D空间感知能力。VLM-IE3D引入了隐式几何标记(IGTs),用于捕获输入视频中的高级几何先验,以及互补的显式几何标记(EGTs),用于编码从重建3D属性中获取的详细几何结构。在此基础上,VLM-IE3D配备了一个3D感知适配器,能够有效融合这两种几何表示与2D视觉线索。这种仅RGB的设计注入了强大的3D归纳偏置,用于细粒度空间理解和推理,无需任何额外的3D输入。大量实验表明,VLM-IE3D在3D视频检测、3D视觉定位、3D密集描述和空间推理等多种3D任务上均取得了一致的优异性能。代码和模型可在 https://github.com/Vegetebird/VLM-IE3D 获取。
原文摘要
Despite rapid progress, most existing vision-language models (VLMs) built from 2D visual inputs often struggle when handling various 3D tasks that require fine-grained spatial understanding and reasoning. To bridge this gap, we present VLM-IE3D, a unified framework that enhances the 3D spatial awareness of VLMs by equipping them with both implicit and explicit 3D geometries learned from RGB videos. Our VLM-IE3D introduces Implicit Geometry Tokens (IGTs) that capture high-level geometric priors from input videos, as well as complementary Explicit Geometry Tokens (EGTs) that encode detailed geometric structures from reconstructed 3D attributes. On top of that, VLM-IE3D comes with a 3D-aware adapter that effectively fuses the two types of geometric representations with 2D visual cues. This RG…
— 自动采集于 2026-07-26
#论文 #arXiv #CV #小凯
