论文概要
研究领域: CV 作者: Lukas Knobel, Andrew Zisserman, Yuki M. Asano 发布时间: 2026-07-25 arXiv: 2507.20472
中文摘要
理解视频中的运动是视觉学习的一个基本挑战,因为帧与帧之间的变化纠缠了两种动力学来源:相机运动和物体运动。这种分解在表示学习中仍未得到充分探索,部分原因是这些因素在自然视频中紧密耦合,难以单独监督。然而恢复它对于学习将有意义物体动力学与相机引起的变异分离的稳健运动表示很重要。我们研究了这种结构化运动表示是否可以从预训练图像视觉Transformer的冻结特征中恢复。我们提出了结构化动力学模型(SDM),它通过未来特征预测显式分离时间变化的主要来源与残余动力学,而不是用单一纠缠的潜变量或非结构化的空间密集过渡标记来表示视频变化。训练结合了真实视频上的自监督学习与合成Kubric数据上场景动力学的弱监督。我们在ProbeMotion上评估SDM,这是一个新的评估套件,涵盖具有相机运动、物体运动和组合动力学的合成和真实视频。SDM优于使用全局CLS或平均池化特征的主干基线,并且在几个探针上与强监督表示如VGGT相比表现良好,尽管使用的监督弱得多。这些结果表明,预训练图像模型可以很容易地被重新用于结构化视频动力学表示,为学习和分析潜在视频动力学提供了有用的归纳偏置。
原文摘要
Understanding motion in video is a fundamental challenge for visual learning, as frame-to-frame change entangles two sources of dynamics: camera motion and object motion. This decomposition has remained underexplored in representation learning, partly because these factors are tightly coupled in natural videos and difficult to supervise separately. Yet recovering it is important for learning robust motion representations that separate meaningful object dynamics from camera-induced variation. We study whether such structured motion representations can be recovered from frozen features of a pretrained image vision transformer. We propose the Structured Dynamics Model (SDM), which explicitly separates the dominant source of temporal change from residual dynamics through future-feature predict…
— 自动采集于 2026-07-26
#论文 #arXiv #CV #小凯
