文章背景与核心概要
奖励建模是强化学习人类反馈(RLHF)、RLAIF以及基于PPO的训练等对齐技术的基础组成部分。然而,其有效性经常受到高质量人类偏好数据稀缺性和异构性的制约。本文介绍了一种名为MARS(Margin and Semantic-Aware Data Augmentation,边界与语义感知数据增强)的自适应框架,旨在优化低资源环境下的奖励建模。
MARS框架引入了两项核心创新:一是自适应分配(Adaptive Allocation),优先对低边界的偏好对进行数据增强,将计算资源集中在最需要的地方;二是语义精炼(Semantic Refinement),在生成合成训练样本之前,利用基于语义距离的精炼来增强“已选”与“拒绝”响应之间的对比度。在三个偏好数据集和两个奖励模型骨干网络上的实证结果表明,MARS在性能上始终优于标准的均匀增强、WoN以及AdaBoost风格的基线,从而带来了更高的RewardBench得分和更强的对齐胜率。
MARS: Margin and Semantic-Aware Data Augmentation for Reward Modeling
Authors: Payel Bhattacharjee, Osvaldo Simeone, Ravi Tandon
arXiv ID: 2602.17658
Primary Subject: Machine Learning (cs.LG)
Last Revised: August 5, 2026
Authors: Payel Bhattacharjee, Osvaldo Simeone, Ravi Tandon
arXiv ID: 2602.17658
Primary Subject: Machine Learning (cs.LG)
Last Revised: August 5, 2026
Summary
奖励建模是强化学习人类反馈(RLHF)、RLAIF以及基于PPO训练等对齐技术的基础组成部分。然而,其有效性经常受到高质量人类偏好数据稀缺性和异构性的阻碍。
Reward modeling is a foundational component of alignment techniques such as RLHF (Reinforcement Learning from Human Feedback), RLAIF, and PPO-based training. However, its effectiveness is frequently hampered by the scarcity and heterogeneity of high-quality human preference data.
MARS(边界与语义感知数据增强,Margin and Semantic-Aware Data Augmentation)是一个旨在优化低资源环境中奖励建模的自适应框架。该框架引入了两项关键创新: * 自适应分配: 优先对低边界的偏好对进行数据增强,将计算资源集中在最需要的地方。 * 语义精炼: 利用基于语义距离的精炼,在生成合成训练样本之前增强“已选”与“拒绝”响应之间的对比。
在三个偏好数据集和两个奖励模型骨干网络上的实证结果表明,MARS始终优于标准的均匀增强、WoN以及AdaBoost风格的基线,从而实现更高的RewardBench得分和更高的对齐胜率。
MARS (Margin and Semantic-Aware Data Augmentation) is an adaptive framework designed to optimize reward modeling in low-resource environments. The framework introduces two key innovations: * Adaptive Allocation: It prioritizes augmentation for preference pairs with low margins, focusing computational resources where they are most needed. * Semantic Refinement: It utilizes semantic-distance-based refinement to enhance the contrast between "chosen" and "rejected" responses before generating synthetic training samples.
Empirical results across three preference datasets and two reward-model backbones demonstrate that MARS consistently outperforms standard uniform augmentation, WoN, and AdaBoost-style baselines, leading to improved RewardBench scores and higher alignment win rates.
Key Features
- 受控增强: 根据偏好对的难度动态调整数据生成强度。
- 增强对比: 通过精炼对之间的语义距离,提高合成数据的质量。
- 稳健的性能: 通过下游对齐任务和独立裁判评估进行验证,确认性能提升是该方法固有的,而非模型耦合的产物。
Key Features
- Controlled Augmentation: Dynamically adjusts the intensity of data generation based on the difficulty of the preference pair.
- Enhanced Contrast: Improves the quality of synthetic data by refining the semantic distance between pairs.
- Robust Performance: Validated through downstream alignment tasks and independent-judge evaluations, confirming that performance gains are intrinsic to the method rather than artifacts of model coupling.
Access the Paper
Access the Paper
License & Metadata

License & Metadata
- License: Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International
- DOI: https://doi.org/10.48550/arXiv.2602.17658