文章背景与核心概要
在标准分布内视觉环境下,视动模仿策略(Visuomotor imitation policies)通常表现卓越,但当引入视觉相似的物体或容器时,往往会出现性能崩溃。本文通过条件视觉基础定位(conditional visual grounding)的视角深入研究了这一局限性——即成功控制所需的正确视觉目标会随着操作阶段和复杂任务状态的变化而动态调整。
研究人员以动作分块Transformer(Action Chunking with Transformers, ACT)为基础,系统性地引入了在颜色和形状上具有相似性的干扰物,并将策略失效精准定位在抓取和放置阶段。为了解决这些脆弱性,本研究评估了一系列互补的干预措施,例如干扰物数据增强、阶段相关注意力正则化以及基于外观的视觉提示(visual prompting)。这些方法在不损害空间控制数据的前提下,显著增强了目标选择能力,并在仿真环境和真实的 UR3e 机器人平台上大幅提升了系统的鲁棒性。此外,作者在预训练的视觉-语言-动作策略中也观察到了类似的失效模式与恢复效果,证明了有针对性的视觉基础定位改进可以泛化到不同的策略学习范式中。
What Matters, When? Diagnosing and Improving Conditional Visual Grounding in Visuomotor Imitation Policies
Authors: Vivek Chavan, Pengtao Xie, Yahuan Shi, Oliver Heimann, Kevin Haninger, Jörg Krüger
Subjects: Robotics (cs.RO), Artificial Intelligence (cs.AI), Computer Vision and Pattern Recognition (cs.CV)
arXiv: 2609.05376
Submitted: 4 September 2026 (Accepted as an extended abstract at the DexHAND Workshop, ECCV 2026)
Executive Summary
视动模仿策略在标准的分布内视觉环境中通常表现优异,然而当引入视觉相似的物体或容器时却容易受挫。本文通过条件视觉基础定位的视角来研究这一局限性:即在不同的操作阶段和复杂的任务状态下,正确的视觉目标会发生动态变化。
通过使用动作分块Transformer(ACT),研究人员系统性地引入了颜色和形状相似的干扰物体,并将失效精准定位在抓取和放置阶段。为了应对这些脆弱性,本研究评估了多项互补的干预措施——例如干扰物增强、阶段相关注意力正则化以及基于外观的视觉提示——这些方法在不损害空间控制数据的前提下,成功增强了目标选择能力。这些改进极大地提升了仿真环境以及真实 UR3e 机器人平台上的鲁棒性。此外,作者展示了预训练视觉-语言-动作策略中的相似失效模式与恢复效果,证明了针对性的视觉基础定位改进能够跨越不同的策略学习范式实现通用。
Visuomotor imitation policies frequently excel under standard in-distribution visual environments, yet stumble when visually similar objects or receptacles are introduced. This paper investigates this limitation through the lens of conditional visual grounding: the reality that the correct visual target changes dynamically across different manipulation phases and complex task states.
Using Action Chunking with Transformers (ACT), the researchers systematically introduced color- and shape-similar distractor objects and localized failures specifically to the picking and placement stages. To address these vulnerabilities, the study evaluates complementary interventions—such as distractor augmentation, phase-dependent attention regularization, and appearance-based visual prompting—which successfully bolster target selection without compromising spatial control data. These enhancements drastically improve robustness both in simulation and on a physical UR3e robotic setup. Furthermore, the authors demonstrate similar failure patterns and recovery in pretrained vision-language-action policies, proving that targeted visual grounding improvements generalize across distinct policy-learning paradigms.
Abstract
视动模仿策略能够在分布内的视觉条件下实现高性能,但在引入视觉相似的物体或容器时则会失效。我们将这种行为作为条件视觉基础定位的问题来研究:成功控制所需的视觉目标会随着操作阶段的变化而改变,在更复杂的任务中还会随着观测到的任务状态而改变。
通过使用动作分块Transformer(ACT),我们系统性地引入了具有受控颜色和形状相似性的干扰物体和容器,并将失效情况定位在抓取和放置阶段。我们发现,干扰物敏感性对视觉相似性的类型和操作阶段都具有特异性。
在此诊断的引导下,我们评估了干扰物增强、阶段相关注意力正则化以及基于外观的视觉提示作为改进目标选择的互补干预手段,同时保留了控制所需的空间信息。这些干预措施大大提高了仿真以及真实 UR3e 机器人上的鲁棒性。我们进一步研究了预训练的视觉-语言-动作策略在状态条件仪器处理任务中的相同失效模式,其中医疗仪器的观测状态决定了正确的目标位置。总之,这些结果表明,即使底层的操作技能保持完整,视觉干扰物也可能导致错误的对象或目的地选择,而显式地改善目标选择可以显著恢复不同视动策略学习机制下的性能。
Visuomotor imitation policies can achieve high performance under in-distribution visual conditions yet fail when visually similar objects or receptacles are introduced. We study this behavior as a problem of conditional visual grounding: the visual target required for successful control changes with the manipulation phase and, in more complex tasks, with the observed task state.
Using Action Chunking with Transformers (ACT), we systematically introduce distractor objects and receptacles with controlled color and shape similarity and localize failures to picking and placement. We find that distractor sensitivity is specific to both the type of visual similarity and the manipulation stage.
Guided by this diagnosis, we evaluate distractor augmentation, phase-dependent attention regularization, and appearance-based visual prompting as complementary interventions for improving target selection while preserving spatial information required for control. These interventions substantially improve robustness in simulation and on a physical UR3e. We further examine the same failure pattern in a pretrained vision-language-action policy on a state-conditioned instrument-handling task, where the observed state of a medical instrument determines the correct destination. Together, the results show that visual distractors can cause incorrect object or destination selection even when the underlying manipulation skill remains intact, and that explicitly improving target selection can substantially recover performance across distinct visuomotor policy-learning regimes.
Links & Resources
- 全文访问: 查看 PDF | HTML 版本 | TeX 源码
- 持久标识符: arXiv:2609.05376 | DOI: 10.48550/arXiv.2609.05376
- Full-Text Access: View PDF | HTML Version | TeX Source
- Persistent Identifiers: arXiv:2609.05376 | DOI: 10.48550/arXiv.2609.05376