跳转至

AeroDPO:通过高保真感知与自动化偏好优化释放轻量化无人机导航潜能

文章背景与核心概要

无人机视觉语言导航(UAV-VLN)要求机器人在复杂的三维环境中具备快速、反应式的控制能力。尽管极简的端到端范式展现出巨大的潜力,但传统上它们严重依赖拥有数十亿参数的大型语言模型,这给现实世界的边缘设备部署带来了极高且难以承受的延迟。

AeroDPO 挑战了这种重参数化的范式,证明了感知质量在根本上重于语言推理能力。该研究的核心要点与创新包括:轻量化效率,通过将紧凑的 2B 模型与高保真视觉输入相结合,成功达到了远超 7B 基准模型的整体成功率;克服行为克隆(BC)的缺陷,纯粹的 BC 缺乏明确的负反馈,导致智能体在处理空间约束时表现困难,并在分布外(OOD)场景中表现出极高的碰撞率;自动化数据飞轮,AeroDPO 引入了一个由确定性物理仿真状态回滚驱动的零成本自动化直接偏好优化(DPO)流水线,当发生碰撞时,系统能自动回溯环境、提取因果推理错误作为被拒绝的动作、应用解耦的特权干预来合成防碰撞的优选动作,并使用离线视觉语言检查器滤除视觉模糊性;最先进的性能,通过将 2B 模型与此自动化流水线相结合,AeroDPO 将未映射场景中的成功率提升至 49.16%,同时大幅降低了碰撞率。


作者: Peng Xu, Chengcheng Wang, Shaohua Wan
学科: 机器人学 (cs.RO);人工智能 (cs.AI);计算机视觉与模式识别 (cs.CV)
arXiv: 2608.07557 [cs.RO] | 提交时间: 2026年8月2日


📌 摘要

Vision-Language Navigation for Unmanned Aerial Vehicles (UAV-VLN) requires rapid, reactive control within complex 3D environments. While minimalist end-to-end paradigms show strong potential, they traditionally rely on massive language models with billions of parameters—introducing prohibitive latency for edge deployment in the real world.

视觉语言无人机导航(UAV-VLN)要求在复杂的三维环境中进行快速、反应式的控制。虽然极简的端到端范式显示出强劲的潜力,但它们传统上依赖于具有数十亿参数的大型语言模型,这为现实世界中的边缘部署带来了过高的延迟。

AeroDPO challenges this parameter-heavy paradigm by demonstrating that perception quality fundamentally outweighs language reasoning capacity. The key takeaways and innovations include: * Lightweight Efficiency: A compact 2B model paired with high-fidelity visual inputs successfully matches the overall success rates of much larger 7B baselines. * Overcoming Behavior Cloning (BC) Flaws: Pure BC lacks explicit negative feedback, causing agents to struggle with spatial constraints and exhibit high collision rates in out-of-distribution (OOD) scenarios. * Automated Data Flywheel: AeroDPO introduces a zero-cost automated Direct Preference Optimization pipeline driven by deterministic physical simulation state rollbacks. When a collision occurs, the system: 1. Autonomously rewinds the environment to extract causal reasoning errors as rejected actions. 2. Applies decoupled privileged interventions to synthesize collision-avoidance preferred maneuvers. 3. Uses an offline vision-language inspector to filter out visual ambiguities. * State-of-the-Art Performance: By coupling the 2B model with this automated pipeline, AeroDPO increases success rates to 49.16% in unmapped scenarios while drastically cutting down collision rates.

AeroDPO 通过证明感知质量在根本上超越了语言推理能力,向这种重参数范式提出了挑战。其核心收获与创新包括: * 轻量化效率: 一个紧凑的 2B 模型与高保真视觉输入相结合,成功匹配了更大规模的 7B 基准模型的整体成功率。 * 克服行为克隆(BC)缺陷: 纯粹的 BC 缺乏明确的负反馈,导致智能体在空间约束方面遇到困难,并在分布外(OOD)场景中表现出高碰撞率。 * 自动化数据飞轮: AeroDPO 引入了一个由确定性物理仿真状态回滚驱动的零成本自动化直接偏好优化流水线。当发生碰撞时,系统会: 1. 自主回溯环境,将因果推理错误提取为拒绝动作(rejected actions)。 2. 应用解耦的特权干预来合成避碰优选动作(preferred maneuvers)。 3. 使用离线视觉语言检查器过滤掉视觉模糊性。 * 最先进的性能: 通过将 2B 模型与该自动化流水线相结合,AeroDPO 在未映射场景中的成功率提高到了 49.16%,同时显着降低了碰撞率。


📋 论文元数据

详情 信息
标识符 arXiv:2608.07557
主要学科 机器人学 (cs.RO)
引用格式 arXiv:2608.07557 [cs.RO]
DOI 10.48550/arXiv.2608.07557
文档统计 7页,3幅图,4张表

🔗 全文与访问链接