论文概要
研究领域: CV 作者: Shing Ho J. Lin, Wenzhao Zheng, Dong Zhuo, Yuqi Wu, Jie Zhou, Jiwen Lu 发布时间: 2026-07-24 arXiv: 2607.22534
中文摘要
几何基础模型(GFMs)大幅推进了单目3D重建,但将这一能力扩展到4D动态理解仍然是一个根本性挑战。大多数现有运动感知方法(如稀疏跟踪、稠密点级光流)将运动视为独立的点级位移,忽略了物理运动的结构化特性。然而,真实世界物体通常遵循刚体运动学,因此点通常是集体运动而非孤立运动。运动本身具有几何结构:物理物体经历由SE(3)控制的一组刚体变换,而非无结构的点级位移。基于这一洞察,本文提出SM4RT,一种用于端到端3D重建和结构化运动感知的结构化运动4D重建Transformer。SM4RT引入运动结构来表示场景动态,其中场景运动被分解为一组紧凑的运动基,每个基表示为SE(3)中6D旋量的时间序列。然后通过在这些基上的稀疏、时间共享的逐像素分配权重恢复稠密场景运动,确保同一物体上的点共享共同的刚体运动轨迹。SM4RT引入并行的运动几何编码器和解码器,从单目RGB视频单次前向传播中联合推断3D几何、世界坐标运动和场景运动学结构。SM4RT在保持场景运动几何结构的同时实现了强大的运动重建性能。
原文摘要
Geometry Foundation Models (GFMs) have substantially advanced monocular 3D reconstruction, yet extending this capability to 4D dynamic understanding remains a fundamental challenge. Most existing motion perception methods (e.g., sparse tracking, dense point-wise flow) treat motion as independent point-wise displacements, ignoring the structured nature of physical motion. However, real-world objects usually obey rigid-body kinematics, and points thus usually move collectively, not in isolation. Motion itself possesses geometric structure: physical objects undergo a set of rigid-body transformations governed by SE(3), rather than unstructured point-wise displacements. Building on this insight, we propose SM4RT, a Structured Motion 4D Reconstruction Transformer for end-to-end 3D reconstructio…
— 自动采集于 2026-07-28
#论文 #arXiv #CV #小凯
