Co-RL:多智能体强化学习中多元群组涌现的无监督推理
文章背景与核心概要
强化学习(RL)已成为增强语言模型和视觉语言模型(VLM)推理能力的核心支柱。然而,传统方法严重依赖成本高昂的真实标签监督奖励。尽管“自奖励(self-rewarding)”模型试图通过生成自身的反馈来绕过这一限制,但它们往往会遭受训练崩塌、响应同质化以及现有偏见自我强化等问题的困扰。
Co-RL 是一个旨在通过协同多智能体训练来解决这些局限性的全新框架。通过利用由解耦且参数独立的模型组成的多元群组,Co-RL 从同行评估(peer evaluation)中获取奖励信号,而非依赖自我反思。这种通过改变模型系列、规模以及对训练样本进行改写而实现的多元性,有效地打破了自我强化的反馈循环。其实验结果证明,这是一个强大且无需标签的训练范式,能够在文本和多模态基准测试中持续超越基础模型,并媲美监督学习的性能。
摘要
Authors: Yunhao Yang, Yuexin Bian, Yunjie Tian, Di Fu, Tianjin Huang, Yuanyuan Shi, Ziang Xiao, Nuno Vasconcelos, Yijiang Li
Date: August 18, 2026
Subject: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computer Vision (cs.CV)
Links: View PDF | Code RepositoryReinforcement Learning (RL) has become a cornerstone for enhancing reasoning in language and vision-language models. However, traditional approaches rely heavily on costly, ground-truth supervised rewards. While "self-rewarding" models attempt to bypass this by generating their own feedback, they often suffer from training collapse, response homogenization, and the reinforcement of existing biases.
Co-RL is a novel framework that addresses these limitations through cooperative multi-agent training. By utilizing a diverse cohort of decoupled, parameter-independent models, Co-RL derives reward signals from peer evaluation rather than self-reflection. This diversity—achieved through varying model families, sizes, and rephrased training samples—effectively breaks self-reinforcing feedback loops. The result is a robust, label-free training paradigm that consistently outperforms base models and matches supervised performance across both text and multimodal benchmarks.
核心贡献
Key Contributions
- 解耦的多智能体训练: 引入了一个多个模型相互优化的框架,消除了对真实标签(ground-truth labels)的需求。
- 缓解训练崩塌: 证明了增加群组多元性可以减少相关错误(correlated errors),从而防止自奖励系统中常见的同质化现象。
- 性能提升:
- 大语言模型(LLMs): 在七个纯文本基准测试中平均提升了 3.0–8.6%。
- 视觉语言模型(VLMs): 在四个多模态基准测试中平均提升了 2.3–7.2%。
- 可扩展性: 证明了无需人工干预,无监督推理可以通过基于同行的奖励结构自然涌现。
- Decoupled Multi-Agent Training: Introduces a framework where multiple models optimize each other, eliminating the need for ground-truth labels.
- Mitigation of Training Collapse: Demonstrates that increasing cohort diversity reduces correlated errors, preventing the homogenization common in self-rewarding systems.
- Performance Gains:
- LLMs: Achieved average gains of 3.0–8.6% across seven text-only benchmarks.
- VLMs: Achieved average gains of 2.3–7.2% across four multimodal benchmarks.
- Scalability: Proves that unsupervised reasoning can emerge naturally from peer-based reward structures without human intervention.
元数据
Metadata
- arXiv ID: 2608.17253
- DOI: 10.48550/arXiv.2608.17253
- 格式: 30页,5张图表,11张表格
- arXiv ID: 2608.17253
- DOI: 10.48550/arXiv.2608.17253
- Format: 30 pages, 5 figures, 11 tables