论文概要
研究领域: CV 作者: Simon Khan, Laurent Gajny, Jennyfer Lecompte, Sébastien Laporte 发布时间: 2026-09-09 arXiv: 2609.10498
中文摘要
从单目体育转播中恢复3D人体姿态仍然具有挑战性,尤其是当球员需要在共享度量世界坐标系中定位而不仅仅是相对于自身身体重建时。本文提出 Field Converter,一种用于从校准足球转播中进行世界 grounded 3D球员姿态估计的几何初始化时间残差框架。该方法首先使用相机和球场几何通过射线-地面交点初始化球员根节点,然后从姿态、图像、相机和几何线索预测时间残差校正。在匹配不相关的评估序列上,残差细化将根节点误差从仅几何的49cm降低到帧级MLP的14cm和TCN的10cm,Transformer达到可比的11cm。结果世界空间MPJPE达到13.2cm。
原文摘要
Recovering 3D human pose from monocular sports broadcasts remains challenging when players must be localized in a shared metric world coordinate system rather than only reconstructed relative to their own body. We introduce Field Converter, a geometry-initialized temporal residual framework for world-grounded 3D player pose estimation from calibrated soccer broadcasts. Our method first uses camera and pitch geometry to initialize the player root through ray-ground intersection, then predicts a temporal residual correction from pose, image, camera, and geometric cues. On match-disjoint evaluation sequences, residual refinement reduces root error from 49cm with geometry alone to 14cm with a frame-wise MLP and 10cm with a TCN, while a Transformer achieves a comparable 11cm. The resulting world-space MPJPE reaches 13.2cm, and ablations show that residual prediction clearly outperforms direct global-root regression while temporal context matters more than the specific temporal backbone. Failure analysis further identifies airborne motion as the main limitation of the ground-based geometric initialization.
— 自动采集于 2026-09-11
#论文 #arXiv #CV #小凯
