尽管许多开放式设置依赖于基于规则的奖励,但具有可验证奖励的强化学习在数学和编码等领域实现了强大的训练后收益,。我们研究了基于标题的 RL, 中的奖励黑客行为,其中策略针对训练验证者进行了优化,但针对由三名前沿法官组成的跨家庭小组进行了评估,,减少了对任何单个评估者的依赖。我们的框架将分歧的两个来源分开:验证者失败,,其中训练验证者相信参考验证者拒绝的评价标准,和评价标准设计限制,,其中即使是基于评价标准的强大验证者也倾向于无评价的法官总体评价更差的反应。在医学和科学领域,弱验证者产生巨大的代理奖励收益,但不会转移到参考验证者;利用随着训练而增长,并集中于反复出现的失败,例如复合标准的部分满足,将隐式内容视为显式,和不精确的主题匹配。更强的验证者可以显着减少,,但并不能消除, 验证者的利用。我们还引入了自我内在化差距,,这是一种基于策略日志概率,的无验证者诊断,它跟踪参考验证者质量,,检测使用弱验证者训练的策略何时停止改进。最后,在我们的设置中,当评分标准未指定重要的故障模式时,更强的验证并不能阻止奖励黑客:基于评分标准的验证者更喜欢RL检查点,,而无评分标准的法官更喜欢基本模型。这些分歧与集中在完整性和基于存在的标准, 上的收益相一致,同时事实正确性, 简洁性, 相关性, 和整体质量的下降。 , 这些结果表明,更强的验证可以减少奖励黑客,,但其本身并不能确保标题收益与更广泛的质量收益相对应。

Reinforcement learning with verifiable rewards has enabled strong post-training gains in domains such as math and coding, though many open-ended settings rely on rubric-based rewards. We study reward hacking in rubric-based RL, where a policy is optimized against a training verifier but evaluated against a cross-family panel of three frontier judges, reducing dependence on any single evaluator. Our framework separates two sources of divergence: verifier failure, where the training verifier credits rubric criteria that reference verifiers reject, and rubric-design limitations, where even strong rubric-based verifiers favor responses that rubric-free judges rate worse overall. Across medical and science domains, weak verifiers produce large proxy-reward gains that do not transfer to the reference verifiers; exploitation grows over training and concentrates in recurring failures such as partial satisfaction of compound criteria, treating implicit content as explicit, and imprecise topical matching. Stronger verifiers substantially reduce, but do not eliminate, verifier exploitation. We also introduce a self-internalization gap, a verifier-free diagnostic based on policy log-probabilities, which tracks reference-verifier quality, detecting when the policy trained using the weak verifier stops improving. Finally, in our setting, stronger verification does not prevent reward hacking when the rubric leaves important failure modes unspecified: rubric-based verifiers prefer the RL checkpoint, while rubric-free judges prefer the base model. These disagreements coincide with gains concentrated in completeness and presence-based criteria, alongside declines in factual correctness, conciseness, relevance, and overall quality. Together, these results suggest that stronger verification reduces reward hacking, but does not by itself ensure that rubric gains correspond to broader quality gains.

科目:人工智能(cs.AI)

Subjects: Artificial Intelligence (cs.AI)