通过算法与超参数的SHAP分析增强机器人强化学习的泛化能力
文章背景与核心概要
本文提出了一种可解释的机器学习框架,旨在解决机器人强化学习(RL)中长期存在的关键泛化鸿沟问题。尽管强化学习的性能在很大程度上受到算法和超参数选择的影响,但以往的研究缺乏一种量化方法来解构这些影响。通过应用 SHapley Additive exPlanations (SHAP),作者评估了不同机器人环境中配置的影响,建立了将 Shapley 值与泛化能力相连接的理论基础,并引入了一种 SHAP 引导的配置选择策略,为真实世界中的机器人部署提供了可操作的指导。
Summary
This paper presents an explainable machine learning framework to address the critical generalizability gap in Reinforcement Learning (RL) for robotics. While RL performance is heavily influenced by algorithm and hyperparameter choices, prior research has lacked a quantitative approach to decompose these impacts. By applying SHapley Additive exPlanations (SHAP), the authors evaluate configuration impacts across diverse robotic environments, establish a theoretical foundation linking Shapley values to generalizability, and introduce a SHAP-guided configuration selection strategy that offers actionable guidance for real-world robotic deployments.
Document Metadata
Metadata Field Details arXiv ID arXiv:2605.02867[cs.LG]Authors Lingxiao Kong, Cong Yang, Oya Deniz Beyan, Zeyd Boukhers Primary Subject Machine Learning ( cs.LG)Secondary Subjects Artificial Intelligence ( cs.AI), Robotics (cs.RO)Accepted Venue International Conference on Pattern Recognition (ICPR) 2026 License Creative Commons Attribution-ShareAlike 4.0 International
Abstract
尽管强化学习(RL)取得了显著进展,但模型性能对算法和超参数配置仍然高度敏感,同时跨环境的泛化鸿沟也使真实世界的部署变得复杂。
Despite significant advances in Reinforcement Learning (RL), model performance remains highly sensitive to algorithm and hyperparameter configurations, while generalization gaps across environments complicate real-world deployment.
尽管先前的工作已经研究了强化学习的泛化能力,但特定配置对泛化鸿沟的相对贡献尚未得到量化解构,也未被系统地利用于配置选择中。为了解决这一局限性,我们提出了一种可解释的框架,该框架通过使用 SHapley Additive exPlanations (SHAP) 来量化配置影响,从而评估不同机器人环境中的 RL 性能。
Although prior work has studied RL generalization, the relative contribution of specific configurations to the generalization gap has not been quantitatively decomposed and systematically leveraged for configuration selection. To address this limitation, we propose an explainable framework that evaluates RL performance across robotic environments using SHapley Additive exPlanations (SHAP) to quantify configuration impacts.
核心贡献包括: * 建立了将 Shapley 值与泛化能力相联系的理论基础。 * 对配置影响模式进行了实证分析。 * 引入了 SHAP 引导的配置选择 来增强泛化能力。
Key contributions include: * Establishing a theoretical foundation connecting Shapley values to generalizability. * Empirically analyzing configuration impact patterns. * Introducing SHAP-guided configuration selection to enhance generalization.
我们的结果揭示了不同算法和超参数的独特模式,并在各种任务和环境中保持了稳定的配置影响。通过将这些见解应用于配置选择,我们提高了 RL 的泛化能力,并为从业者提供了可操作的指导。
Our results reveal distinct patterns across algorithms and hyperparameters, with consistent configuration impacts across diverse tasks and environments. By applying these insights to configuration selection, we achieve improved RL generalizability and provide actionable guidance for practitioners.
Submission History
- [v1] 2026年5月4日 星期一 17:41:04 UTC (1,675 KB)
- [v2] 2026年6月22日 星期一 14:29:50 UTC (1,677 KB)
- [v3] 2026年8月25日 星期二 13:57:38 UTC (1,677 KB) — 此版本
- [v1] Mon, 4 May 2026 17:41:04 UTC (1,675 KB)
- [v2] Mon, 22 Jun 2026 14:29:50 UTC (1,677 KB)
- [v3] Tue, 25 Aug 2026 13:57:38 UTC (1,677 KB) — This version
