跳转至

文章背景与核心概要

在自动驾驶领域,利用强化学习(RL)来保证安全性一直受到现实世界交通“长尾效应”的阻碍,即安全关键场景在日常中极为罕见。现有的方法通常独立于智能体的当前策略来生成对抗性场景,导致训练效率低下。

本文引入了一种名为威胁引导的策略感知场景扰动(Threat-guided Policy-aware Scene Perturbation, TPSP)的新方法。TPSP 利用策略感知场景编码器来识别智能体行为与其环境之间的交互。它不进行统一的场景修改,而是有选择地对关键物体进行扰动。通过评估原始场景与扰动场景之间的“威胁水平差异”,该框架能够生成高价值、安全关键的训练数据,在有限的交互预算下显著提升安全性学习效率。


Threat-guided Policy-aware Scene Perturbation for Safe Autonomous Driving with Online Reinforcement Learning

Threat-guided Policy-aware Scene Perturbation for Safe Autonomous Driving with Online Reinforcement Learning

Authors: Xincong Hu, Lei Ou, Maosen Li, Jingtao Zhang, Liguo Hou, Zongzhang Zhang
Affiliations: Nanjing University, Yinwang Intelligent Technology Co., Ltd
Date: August 11, 2026
arXiv ID: 2608.10403

Authors: Xincong Hu, Lei Ou, Maosen Li, Jingtao Zhang, Liguo Hou, Zongzhang Zhang
Affiliations: Nanjing University, Yinwang Intelligent Technology Co., Ltd
Date: August 11, 2026
arXiv ID: 2608.10403


Summary

Summary

确保通过强化学习(RL)实现自动驾驶的安全,受到现实世界交通“长尾”特性的阻碍,在这一特性中,安全关键场景非常罕见。现有方法通常独立于智能体当前的策略来生成对抗场景,从而导致训练效率低下。

Ensuring safety in autonomous driving via Reinforcement Learning (RL) is hindered by the "long-tailed" nature of real-world traffic, where safety-critical scenarios are rare. Existing methods often generate adversarial scenes independently of the agent's current policy, leading to inefficient training.

本文介绍了威胁引导的策略感知场景扰动(TPSP)。TPSP 利用策略感知场景编码器来识别智能体行为与环境之间的交互。它不进行统一的场景修改,而是有选择地扰动关键物体。通过评估原始场景和扰动场景之间的“威胁水平差异”,该框架生成高价值、安全关键的训练数据,在有限的交互预算下显着提高安全性学习效率。

This paper introduces Threat-guided Policy-aware Scene Perturbation (TPSP). TPSP utilizes a policy-aware scene encoder to identify the interaction between the agent's behavior and its environment. Instead of uniform scene modification, it selectively perturbs critical objects. By evaluating the "threat-level difference" between original and perturbed scenarios, the framework generates high-value, safety-critical training data, significantly improving safety learning efficiency under limited interaction budgets.


Key Contributions

Key Contributions

  • 策略感知场景编码: 一种新颖的机制,可捕获自动驾驶策略与其周围环境之间的动态关系。
  • 定向扰动: 与随机或全局场景修改不同,TPSP 专注于扰动特定的关键对象,以最大化训练经验的信息价值。
  • 威胁引导优化: 一种根据威胁水平方差优先生成场景的策略,确保智能体接触到最相关的安全关键挑战。
  • 实证验证: 该方法在 NAVSIM v2 基准测试中进行了验证,成功处理了大约 400 万公里的模拟驾驶数据,证明了与传统基线相比优越的安全性能。
  • Policy-Aware Scene Encoding: A novel mechanism that captures the dynamic relationship between the autonomous driving policy and its surrounding environment.
  • Targeted Perturbation: Unlike random or global scene modifications, TPSP focuses on perturbing specific critical objects to maximize the informative value of the training experience.
  • Threat-Guided Optimization: A strategy that prioritizes the generation of scenes based on the threat-level variance, ensuring the agent is exposed to the most relevant safety-critical challenges.
  • Empirical Validation: Demonstrated on the NAVSIM v2 benchmark, the method successfully processed approximately 4 million kilometers of simulated driving data, proving superior safety performance compared to traditional baselines.

Metadata

Metadata

Category Details
Primary Subject Artificial Intelligence (cs.AI)
Comments 11 pages, 5 figures
DOI 10.48550/arXiv.2608.10403
Category Details
Primary Subject Artificial Intelligence (cs.AI)
Comments 11 pages, 5 figures
DOI 10.48550/arXiv.2608.10403