论文概要
研究领域: CV 作者: Jiacong Xu, Hanwen Jiang, Zhixin Shu, Kalyan Sunkavalli, Vishal M. Patel, Yiqun Mei 发布时间: 2026-07-28 arXiv: 2607.26037
中文摘要
我们提出了Wonder,一个通用的视频世界模型,用于实时、相机可控的世界探索。给定一张图像或条件视频,Wonder构建一个可交互的世界,用户可以通过移动相机进行交互式导航,发现未见的区域,并在实时和长期范围内重新访问先前观测的区域。实现这一能力需要控制方法、记忆机制和训练策略的系统级协同设计。我们引入了一种新颖的相机条件方法,使用密集坐标场,其渲染提供空间对齐的运动和方向线索,使模型能直接将相机运动解释为视觉证据。为了支持在增长的生成上下文中快速精确的记忆检索,我们提出了一种高效的稀疏注意力记忆机制,使模型在推理时能选择性地关注一小部分相关上下文token,而不管实际上下文长度。我们进一步开发了多种技术来修正自强制风格蒸馏管道,提高学生模型尊重控制信号的能力,同时保持教师模型的多样生成模式和长期记忆。这些组件共同使Wonder能够以16 FPS合成多样化的分钟级视频,同时在长程展开中保持连贯的几何、外观和动态。除了图像到视频生成,Wonder自然支持视频条件生成,允许现有动态场景被实时重拍。
原文摘要
We present Wonder, a general-purpose video world model for real-time, camera-controllable world exploration. Given an image or a conditional video, Wonder constructs a playable world where users can navigate interactively by moving the camera, discovering unseen regions, and revisiting previously observed areas in real time and over a long-term horizon. Achieving this capability requires a system-level co-design of control method, memory mechanism, and training strategy. We introduce a novel camera conditioning with a dense coordinate field whose renderings provide spatially aligned motion and orientation cues, allowing the model to interpret camera motion directly as visual evidence. To support fast and precise memory retrieval over a growing generation context, we propose an efficient sp…
— 自动采集于 2026-07-30
#论文 #arXiv #CV #小凯
