跳转至

强化学习引导的演化策略优化用于偏好可调的异构敏捷对地观测卫星调度

文章背景与核心概要

随着航天技术的快速发展,敏捷对地观测卫星(AEOS)在对地观测任务中发挥着越来越重要的作用。然而,由于卫星平台具有异构性(即能力和资源消耗各不相同),且必须满足复杂的约束条件(如星载可见窗口、姿态机动要求、能量消耗和机载存储限制),实现统一的建模与优化面临着巨大的挑战。

为了解决这一难题,本文提出了一种新颖的演化策略优化框架,主要包含建模层和优化层。建模层通过结合基于分配的间接编码与基于解码器的等效成本评估,将任务收益、节能情况和负载平衡转化为统一、可解释的标量效用;优化层则解耦了调度解码、基于种群的搜索以及在线动作-评论家(Actor-Critic)算子控制。

在此基础上,作者设计了 RLOSMEA(强化学习辅助的算子选择模因演化算法)。该算法旨在严格的函数评估预算下,平衡全局探索、可行性恢复和局部求精。实验结果表明,RLOSMEA 在整体加权效用和收敛稳定性方面均优于主流的元启发式基线方法。


📋 Summary (摘要翻译与原文)

This paper introduces a novel framework for scheduling heterogeneous Agile Earth Observation Satellites (AEOS) that must satisfy complex constraints—such as satellite-dependent visibility windows, attitude maneuvering requirements, energy consumption, and onboard storage limits. Because satellites vary significantly in their capabilities and resource consumption, unified modeling and optimization remain challenging.

To overcome this, the authors propose an evolutionary policy optimization framework featuring: 1. A Modeling Layer: Combines assignment-based indirect encoding with decoder-based equivalent-cost evaluation. This translates task gains, energy savings, and load balances into a unified, interpretable scalar utility while preserving platform-specific constraints. 2. An Optimization Layer: Decouples schedule decoding, population-based search, and online actor-critic operator control. Instead of building schedules directly, reinforcement learning selects high-level search operators. 3. The RLOSMEA Algorithm: A reinforcement-learning-assisted operator-selection memetic evolutionary algorithm designed to balance global exploration, feasibility recovery, and local refinement under strict function-evaluation budgets.

Experiments demonstrate that RLOSMEA outperforms representative metaheuristic baselines, yielding higher overall weighted utility and more stable convergence.

本文介绍了一种用于调度异构敏捷对地观测卫星(AEOS)的新颖框架。该卫星群必须满足复杂的约束条件,例如卫星相关的可见窗口、姿态机动要求、能量消耗和机载存储限制。由于卫星在能力和资源消耗方面存在显著差异,统一的建模和优化仍然具有挑战性。

为了克服这一问题,作者提出了一种演化策略优化框架,其特点包括: 1. 建模层: 结合了基于分配的间接编码与基于解码器的等效成本评估。在保持平台特定约束的同时,将任务收益、节能和负载平衡转化为统一、可解释的标量效用。 2. 优化层: 解耦了调度解码、基于种群的搜索以及在线动作-评论家算子控制。强化学习不直接构建调度方案,而是选择高级搜索算子。 3. RLOSMEA 算法: 一种强化学习辅助的算子选择模因演化算法,旨在严格的函数评估预算下平衡全局探索、可行性恢复和局部求精。

实验表明,RLOSMEA 优于代表性的元启发式基线,能够产生更高的整体加权效用和更稳定的收敛性。


📄 Metadata & Article Details (元数据与文章详情)

  • arXiv ID: arXiv:2608.24470 [cs.AI]
  • Primary Subject: Artificial Intelligence (cs.AI)
  • Submission Date: 25 August 2026
  • Authors:
  • He Wang
  • Junyu Wu
  • Hui Li
  • Yanjie Song
  • Witold Pedrycz
  • Liang Li
  • Length: 14 pages, 8 figures
  • DOI: 10.48550/arXiv.2608.24470
  • arXiv ID: arXiv:2608.24470 [cs.AI]
  • 主要学科: 人工智能 (cs.AI)
  • 提交日期: 2026年8月25日
  • 作者:
  • He Wang
  • Junyu Wu
  • Hui Li
  • Yanjie Song
  • Witold Pedrycz
  • Liang Li
  • 篇幅页数: 14页,8幅图表
  • DOI: 10.48550/arXiv.2608.24470