[论文] What Matters, When? Diagnosing and Improving Conditional Visual Ground…

## 论文概要 **研究领域**: CV **作者**: Vivek Chavan, Pengtao Xie,...

论文概要

研究领域: CV 作者: Vivek Chavan, Pengtao Xie, Yahuan Shi, Oliver Heimann, Kevin Haninger, Jörg Krüger 发布时间: 2026-09-04 arXiv: 2609.05376

中文摘要

视觉运动模仿策略在分布内视觉条件下可取得高性能,但引入视觉相似物体或容器时失败。本文将此行为作为条件视觉定位问题研究:成功控制所需的视觉目标随操作阶段变化,复杂任务中还随观测任务状态变化。使用Action Chunking with Transformers(ACT),我们系统引入颜色和形状相似度可控的干扰物体和容器,并将失败定位到抓取和放置阶段。发现干扰敏感性对视觉相似类型和操作阶段均特异。基于此诊断,评估干扰增强、阶段依赖注意力正则化和基于外观的视觉提示作为互补干预,在保持控制所需空间信息的同时改善目标选择。这些干预在仿真和物理UR3e上显著提升鲁棒性。我们进一步在预训练视觉-语言-动作策略中检查相同失败模式,观测到的医疗器械状态决定正确目的地。结果表明,即使底层操作技能完好,视觉干扰也可导致错误的物体或目的地选择,显式改善目标选择可在不同视觉运动策略学习范式中大幅恢复性能。

原文摘要

Visuomotor imitation policies can achieve high performance under in-distribution visual conditions yet fail when visually similar objects or receptacles are introduced. We study this behavior as a problem of conditional visual grounding: the visual target required for successful control changes with the manipulation phase and, in more complex tasks, with the observed task state. Using Action Chunking with Transformers (ACT), we systematically introduce distractor objects and receptacles with controlled color and shape similarity and localize failures to picking and placement. We find that distractor sensitivity is specific to both the type of visual similarity and the manipulation stage. Guided by this diagnosis, we evaluate distractor augmentation, phase-dependent attention regularization…

自动采集于 2026-09-09

#论文 #arXiv #CV #小凯

发表回复

人生梦想 - 关注前沿的计算机技术 acejoy.com 🐾 步子哥の博客 🐾 背多分论坛 🐾 借一步网 🐾 智柴网 沪ICP备2024052574号-1