文章背景与核心概要
在机器人自主执行长序列任务时,如何在有限的练习时间和资源(即练习预算)内高效习得所需技能,是机器人学领域的一项核心挑战。为了解决这一问题,本文提出了一种名为“刻意练习(Deliberate Practice, DP)”的新型主动技能学习算法。该算法能够通过估算掌握特定技能所需的时间成本以及这些技能解锁的任务规划所带来的累积回报,智能地决定优先练习哪些技能。
技术核心上,DP 通过构建一个双线性规划模型(bilinear program),利用现成的求解器计算出在预算约束下能够使期望累积回报最大化的最优练习时间分配方案。通过在仿真环境和真实世界的长期操作任务(long-horizon manipulation tasks)中进行验证,该方法被证明能够帮助机器人在有限的练习预算下最优地获取实用策略,从而显著提升长序列任务的规划性能。这一研究为机器人自主技能获取和高效资源分配提供了重要的理论与实验支撑。
Deliberate Practice: Learning Robot Skills under a Budget
arXiv ID: 2608.13415
Subjects: Robotics (cs.RO); Artificial Intelligence (cs.AI)
Date: August 13, 2026
Summary
Deliberate Practice (DP) is a novel active skill learning algorithm designed to help robots autonomously acquire skills when faced with a limited practice budget. The core challenge in long-horizon sequential tasks is determining which skills to prioritize to maximize cumulative reward. DP addresses this by: * Estimating Costs and Rewards: Calculating the time required to master specific skills versus the potential reward gained from the task plans those skills enable. * Budget Optimization: Utilizing a bilinear program to compute a provably budget-optimal allocation of practice time. * Performance: Validated through both simulated and real-world long-horizon manipulation experiments, demonstrating that robots can effectively prioritize learning to improve overall task planning.
Deliberate Practice (DP) 是一种新颖的主动技能学习算法,旨在帮助机器人在面对有限的练习预算时自主获取技能。长序列任务中的核心挑战在于确定优先学习哪些技能以最大化累积回报。DP 通过以下方式解决这一问题: * 估计成本与回报: 计算掌握特定技能所需的时间,与这些技能所解锁的任务计划所获得的潜在回报进行权衡。 * 预算优化: 利用双线性程序来计算出在理论上具有预算最优性的练习时间分配。 * 性能表现: 通过仿真和真实世界的长期操作实验验证,证明机器人能够有效地优先学习,从而改进整体任务规划。
Authors
- Shivam Vats
- Sudarshan Harithas
- Mete Tuluhan Akbulut
- Arvind Raghunathan
- George Konidaris
- Shivam Vats
- Sudarshan Harithas
- Mete Tuluhan Akbulut
- Arvind Raghunathan
- George Konidaris
Abstract
We consider the problem of autonomously learning robot skills under a limited practice budget for sequential tasks. We propose an active skill learning algorithm, Deliberate Practice (DP), that computes a provably budget-optimal allocation—practicing skills that maximize expected cumulative reward while being learnable within the budget. DP estimates both the time needed to master skills and the cumulative reward of the task plans that the skills unlock. Computing a budget-optimal allocation is challenging as it requires reasoning about combinatorially many skill plans over a large practice budget. Our key contribution is a bilinear program that can compute this exactly using off-the-shelf solvers. Through simulated and real-world experiments on long-horizon manipulation tasks, we show that our approach allows robots to optimally use limited practice time to acquire useful policies and improve long-horizon planning.
我们研究了在有限练习预算下为序列任务自主学习机器人技能的问题。我们提出了一种主动技能学习算法——刻意练习 (Deliberate Practice, DP),它能够计算出可证明的预算最优分配方案:练习那些能够在预算内学会并最大化期望累积回报的技能。DP 既能估计掌握技能所需的时间,又能估计这些技能所解锁的任务 plans 的累积回报。计算预算最优分配具有挑战性,因为它需要在大规模练习预算下对组合繁多的技能计划进行推理。我们的核心贡献是一个双线性程序,它可以使用现成的求解器精确地计算出这一方案。通过在长期操作任务上的仿真和真实世界实验,我们表明我们的方法允许机器人最优地利用有限的练习时间来获取有用的策略并改进长期规划。
Access & Resources

访问与资源
Citation & Metadata
- Comments: 16 pages including appendices.
- Cite as: arXiv:2608.13415 [cs.RO]
- External Links: NASA ADS | Google Scholar | Semantic Scholar
引用与元数据
- 备注: 共 16 页,包含附录。
- 引用格式: arXiv:2608.13415 [cs.RO]
- 外部链接: NASA ADS | Google Scholar | Semantic Scholar