论文概要
研究领域: CV 作者: Yuan Yin, Elias Ramzi, Marc Lafon, Valentin Charraut, Victor Bares, Yihong Xu, Éloi Zablocki, Alexandre Boulch, Thibault Buhet, Andrei Bursuc, Matthieu Cord 发布时间: 2026-07-28 arXiv: 2607.26005
中文摘要
模拟中的自我博弈在大规模上产生了稳健的驾驶策略。这种行为已在使用特权向量化观测(如精确位姿和速度,即使对于被遮挡的智能体)的演示中得到证明。这假设感知已解决,并引入了与部署智能体以自我为中心相机视角的部分观测之间的表征差距。常见的修复方法——将特权策略蒸馏到相机输入的学生模型——使学生模仿其自身视角无法证明的决策。相反,我们建立了透视视角自我博弈作为一种实用的训练范式。我们引入了Pictura,一个GPU加速的多智能体驾驶模拟器,在每一步渲染每个智能体的自我中心视角,从源头缓解表征差距。Pictura在单个H100上维持高达500K智能体步/秒(2M图像/秒)。使用Pictura,我们通过普通PPO自我博弈训练Alberti。它是第一个直接从透视图像训练的大规模驾驶自我博弈策略,无需特权观测。训练跨越50B智能体步,约3500万公里驾驶。它接近其特权向量化对应物的驾驶性能,并零样本迁移到Pictura中重新渲染的Waymo Open Motion Dataset布局,在那里它优于特权向量化智能体。
原文摘要
Self-play in simulation produces robust driving policies at scale. Demonstrations of such behavior have been made using privileged vectorized observations such as exact poses and velocities, even for occluded agents. This assumes that perception is solved and introduces a representation gap with the partial observation of a deployed agent driving from the perspective view of egocentric cameras. A common fix, distilling the privileged policy into a camera-input student, leaves the student imitating decisions its own view cannot justify. Instead, we establish perspective-view self-play as a practical training regime. We introduce Pictura, a GPU-accelerated multi-agent driving simulator that renders each agent’s egocentric view at every step, mitigating the representation gap at its source. P…
— 自动采集于 2026-07-30
#论文 #arXiv #CV #小凯
