面向海上监控中异构传感器选择的强化学习
文章背景与核心概要
本文提出了一种针对异构海上传感器网络中单船跟踪的信息增益引导强化学习框架。传统方法在面对复杂的海洋环境时,往往由于计算期望信息增益而面临严重的计算瓶颈,或者需要长期开启所有传感器以维持精度。为此,该研究利用近端策略优化(PPO)智能体来高效选择最优的跟踪传感器,在每个决策时间步仅激活单个传感器,却能达到接近“全时全开”多传感器配置的跟踪性能。
在技术实现上,该框架结合了贝叶斯序贯蒙特卡洛跟踪器,在非线性与非高斯条件下根据噪声测量估计船舶状态,并提供信念状态表示。智能体通过观察信念状态、检测历史、覆盖范围、传感器几何构型以及实际信息增益等特征进行决策。在塞浦路斯圣纳帕码头 CMMI 智能码头测试床的地理参考仿真中进行的测试表明,该方法不仅大幅降低了计算成本,而且在零样本(zero-shot)布局扰动评估中展现出极强的鲁棒性,位置误差方差极小。
执行摘要
本文提出了一种用于异构海上传感器网络中单船跟踪的信息增益引导强化学习框架。为了解决传统期望信息增益评估的计算瓶颈,该方法利用近端策略优化(PPO)智能体来高效选择最优的跟踪传感器。在塞浦路斯圣纳帕码头(Ayya Napa Marina)CMMI 智能码头测试床的地理参考仿真中进行测试,所学到的策略在每个决策时间步仅激活单个传感器的同时,实现了与“常开”多传感器设置相当的跟踪性能。此外,零样本评估表明,在布局扰动下该方法具有强大的稳定性和极小的位置误差方差。
Executive Summary
This paper presents an information-gain-guided reinforcement learning framework for single-vessel tracking within heterogeneous maritime sensor networks. Addressing the computational bottlenecks of traditional expected-information-gain evaluations, the proposed method leverages a Proximal Policy Optimization (PPO) agent to efficiently select optimal tracking sensors. Tested on a georeferenced simulation of the CMMI Smart Marina testbed in Ayia Napa Marina, Cyprus, the learned policy achieves tracking performance comparable to an "always-on" multi-sensor setup while activating only a single sensor per time step. Furthermore, zero-shot evaluations demonstrate robust stability under layout perturbations with minimal positional error variance.
元数据
- arXiv 标识符:
arXiv:2607.22667[cs.AI] - 主要学科: 人工智能 (
cs.AI) - 其他学科: 信息论 (
cs.IT)、机器学习 (cs.LG)、机器人学 (cs.RO)、信号处理 (eess.SP)、系统与控制 (eess.SY) - 会议收录: 已被 IEEE MetroSea 2026 会议 接受(第 13 专场:用于海上态势感知的目标检测、跟踪与传感器融合)
- 作者:
- Andrei Starodubov
- Yaqub Aris Prabowo
- Andreas Hadjipieris
- Roberto Galeazzi
- Ioannis Kyriakides
- 提交时间线: 2026年7月3日提交;2026年9月2日最后修订(第 2 版)。
Metadata
- arXiv Identifier:
arXiv:2607.22667[cs.AI]- Primary Subject: Artificial Intelligence (
cs.AI)- Other Subjects: Information Theory (
cs.IT), Machine Learning (cs.LG), Robotics (cs.RO), Signal Processing (eess.SP), Systems and Control (eess.SY)- Conference Acceptance: Accepted for the IEEE MetroSea 2026 Conference (Special Session 13: Object Detection, Tracking, and Sensor Fusion for Maritime Situational Awareness)
- Authors:
- Andrei Starodubov
- Yaqub Aris Prabowo
- Andreas Hadjipieris
- Roberto Galeazzi
- Ioannis Kyriakides
- Submission Timeline: Submitted on July 3, 2026; last revised on September 2, 2026 (Version 2).
摘要
本文提出了一种用于异构海上传感器网络中单船跟踪的信息增益引导强化学习传感器选择框架。该方法的动机源于信息论传感器管理:学习得到的策略不是激活所有传感器或重复执行计算昂贵的在线期望信息增益评估,而是在每个决策时刻选择一个与跟踪相关的传感器。
贝叶斯序贯蒙特卡洛跟踪器从噪声测量中估计船舶状态,并为非线性与非高斯条件下的调度提供信念表示。近端策略优化(PPO)智能体在塞浦路斯圣纳帕码头 CMMI 智能码头测试床的地理参考仿真中选择五个传感器之一。该策略在测试床实际的五传感器配置上进行训练。
智能体观察信念状态、检测历史、覆盖范围、传感器几何构型以及实现的实际信息增益特征。奖励定义为受可观测性掩码门控的实际信息增益项。最终测试仿真将所提出的框架与以下方法进行了对比: 1. 随机单传感器选择, 2. 同时使用所有传感器的常开(Always-on)传感,以及 3. 先前工作中提出的期望信息增益传感器选择基线。
核心结果
- 效率: 学习到的策略在每个决策时间步仅激活一个传感器的同时,实现了接近常开传感的跟踪性能。
- 计算成本降低: 成功避免了期望信息增益选择所需的计算昂贵的在线熵搜索。
- 鲁棒性: 在未重新训练的情况下,对实际布局配置的十种中度扰动版本进行了额外的零样本评估,结果显示跟踪大体稳定,所有扰动下的位置跟踪误差增幅均保持在 \(1\text{ 米}\) 以下。
Abstract
This paper presents an information-gain-guided reinforcement-learning sensor-selection framework for single-vessel tracking in heterogeneous maritime sensor networks. The proposed approach is motivated by information-theoretic sensor management: instead of activating all sensors or repeatedly performing computationally expensive online expected-information-gain evaluation, a learned policy selects one tracking-relevant sensor at each decision epoch.
A Bayesian sequential Monte Carlo tracker estimates the vessel state from noisy measurements and provides a belief representation for scheduling under nonlinear and non-Gaussian conditions. A Proximal Policy Optimization agent selects one of five sensors in a georeferenced simulation of the CMMI Smart Marina testbed at Ayia Napa Marina, Cyprus. The policy is trained on the testbed's actual five-sensor configuration.
The agent observes belief-state, detection-history, coverage, sensor-geometry, and realized-information-gain features. The reward is defined as a realized-information-gain term gated by an observability mask. Final-test simulations compare the proposed framework with: 1. Random single-sensor selection, 2. Always-on sensing using all sensors simultaneously, and 3. The expected-information-gain sensor-selection baseline proposed in prior work.
Key Results
- Efficiency: The learned policy achieves tracking performance close to always-on sensing while activating only one sensor per decision time step.
- Computational Cost Reduction: It successfully avoids the computationally expensive online entropy search required by expected-information-gain selection.
- Robustness: Additional zero-shot evaluation without retraining on ten moderately perturbed versions of the actual layout configuration showed broadly stable tracking, with any increase in positional tracking error remaining below \(1\text{ meter}\) across all perturbations.
链接与资源
Links and Resources
- Full-Text Access: View PDF | HTML Version | TeX Source
- Citations & Metrics: Google Scholar | Semantic Scholar | NASA ADS