[论文] Unified Video Dense Prediction from Disjoint Data

## 论文概要 **研究领域**: CV **作者**: Yihong Sun, Seoung Wug Oh,...

论文概要

研究领域: CV 作者: Yihong Sun, Seoung Wug Oh, Jiahui Huang 发布时间: 2026-07-25 arXiv: 2507.20486

中文摘要

场景理解需要同时对几何、外观和语义进行预测。然而,现有的任务特定注释分散在不兼容的、领域特定的数据集中。当前的统一系统通过将训练限制在完全共同注释的数据上,或承担伪标签的巨大计算成本来规避这一问题。为缓解这一问题,我们引入了UniD,一个统一的视频模型,联合预测八种密集场景属性——深度、表面法线、语义分割、边界、人体部位、反照率、阴影和材质——所有这些都从离散的、领域特定的数据集中学习。我们提出了一个简单而有效的蒸馏步骤,其中每个任务的专家通过轻量级任务投影器监督统一的主干网络,消除了对注释重叠或伪标签的需求。我们的核心洞见是,预训练扩散模型的强视觉先验足以弥合离散训练源引入的领域差距,实现对训练期间从未见过的场景-任务组合的稳健泛化。UniD在与每个任务专家和多任务基线的竞争中取得了有竞争力的性能,对分布外场景具有强泛化能力,并增强了时间一致性和跨任务一致性。代码和视频结果可在 https://unid-video.github.io/ 获取。

原文摘要

Scene understanding requires simultaneous prediction about geometry, appearance, and semantics. However, existing task-specific annotations are fragmented across incompatible, domain-specific datasets. Current unified systems circumvent this by restricting training to fully co-annotated data, or by incurring the large computational cost of pseudo-labeling. To mitigate this, we introduce UniD, a unified video model that jointly predicts eight dense scene properties-depth, surface normals, semantic segmentation, boundaries, human parts, albedo, shading, and materials-all learned from disjoint, domain-specific datasets. We propose a simple yet effective distillation step in which per-task experts supervise a unified backbone through lightweight task projectors, eliminating the need for annota…

自动采集于 2026-07-26

#论文 #arXiv #CV #小凯

发表回复

人生梦想 - 关注前沿的计算机技术 acejoy.com 🐾 步子哥の博客 🐾 背多分论坛 🐾 借一步网 🐾 智柴网 沪ICP备2024052574号-1