文章背景与核心概要
在自动驾驶领域,利用强化学习(RL)来保证安全性一直受到现实世界交通“长尾效应”的阻碍,即安全关键场景在日常中极为罕见。现有的方法通常独立于智能体的当前策略来生成对抗性场景,导致训练效率低下。
本文引入了一种名为威胁引导的策略感知场景扰动(Threat-guided Policy-aware Scene Perturbation, TPSP)的新方法。TPSP 利用策略感知场景编码器来识别智能体行为与其环境之间的交互。它不进行统一的场景修改,而是有选择地对关键物体进行扰动。通过评估原始场景与扰动场景之间的“威胁水平差异”,该框架能够生成高价值、安全关键的训练数据,在有限的交互预算下显著提升安全性学习效率。
Threat-guided Policy-aware Scene Perturbation for Safe Autonomous Driving with Online Reinforcement Learning
Threat-guided Policy-aware Scene Perturbation for Safe Autonomous Driving with Online Reinforcement Learning
Authors: Xincong Hu, Lei Ou, Maosen Li, Jingtao Zhang, Liguo Hou, Zongzhang Zhang
Affiliations: Nanjing University, Yinwang Intelligent Technology Co., Ltd
Date: August 11, 2026
arXiv ID: 2608.10403
Authors: Xincong Hu, Lei Ou, Maosen Li, Jingtao Zhang, Liguo Hou, Zongzhang Zhang
Affiliations: Nanjing University, Yinwang Intelligent Technology Co., Ltd
Date: August 11, 2026
arXiv ID: 2608.10403
Summary
Summary
确保通过强化学习(RL)实现自动驾驶的安全,受到现实世界交通“长尾”特性的阻碍,在这一特性中,安全关键场景非常罕见。现有方法通常独立于智能体当前的策略来生成对抗场景,从而导致训练效率低下。
Ensuring safety in autonomous driving via Reinforcement Learning (RL) is hindered by the "long-tailed" nature of real-world traffic, where safety-critical scenarios are rare. Existing methods often generate adversarial scenes independently of the agent's current policy, leading to inefficient training.
本文介绍了威胁引导的策略感知场景扰动(TPSP)。TPSP 利用策略感知场景编码器来识别智能体行为与环境之间的交互。它不进行统一的场景修改,而是有选择地扰动关键物体。通过评估原始场景和扰动场景之间的“威胁水平差异”,该框架生成高价值、安全关键的训练数据,在有限的交互预算下显着提高安全性学习效率。
This paper introduces Threat-guided Policy-aware Scene Perturbation (TPSP). TPSP utilizes a policy-aware scene encoder to identify the interaction between the agent's behavior and its environment. Instead of uniform scene modification, it selectively perturbs critical objects. By evaluating the "threat-level difference" between original and perturbed scenarios, the framework generates high-value, safety-critical training data, significantly improving safety learning efficiency under limited interaction budgets.
Key Contributions
Key Contributions
- 策略感知场景编码: 一种新颖的机制,可捕获自动驾驶策略与其周围环境之间的动态关系。
- 定向扰动: 与随机或全局场景修改不同,TPSP 专注于扰动特定的关键对象,以最大化训练经验的信息价值。
- 威胁引导优化: 一种根据威胁水平方差优先生成场景的策略,确保智能体接触到最相关的安全关键挑战。
- 实证验证: 该方法在 NAVSIM v2 基准测试中进行了验证,成功处理了大约 400 万公里的模拟驾驶数据,证明了与传统基线相比优越的安全性能。
- Policy-Aware Scene Encoding: A novel mechanism that captures the dynamic relationship between the autonomous driving policy and its surrounding environment.
- Targeted Perturbation: Unlike random or global scene modifications, TPSP focuses on perturbing specific critical objects to maximize the informative value of the training experience.
- Threat-Guided Optimization: A strategy that prioritizes the generation of scenes based on the threat-level variance, ensuring the agent is exposed to the most relevant safety-critical challenges.
- Empirical Validation: Demonstrated on the NAVSIM v2 benchmark, the method successfully processed approximately 4 million kilometers of simulated driving data, proving superior safety performance compared to traditional baselines.
Metadata
Metadata
| Category | Details |
|---|---|
| Primary Subject | Artificial Intelligence (cs.AI) |
| Comments | 11 pages, 5 figures |
| DOI | 10.48550/arXiv.2608.10403 |
Category Details Primary Subject Artificial Intelligence (cs.AI) Comments 11 pages, 5 figures DOI 10.48550/arXiv.2608.10403
Access Links
Access Links