我们认为,将奖励分解为加权,可验证标准并使用LLM法官对其进行评分提供了部分学分优化信号:,而不是二元结果或单个整体分数,每个响应都根据多个特定于任务的标准进行评分。我们将 \emph{rubric 为基础的强化学习 (RL)}: 正式化为一个框架,在该框架中,策略针对由冻结的 LLM 法官产生的结构化, 多标准奖励进行优化,该策略从未见过辅助接地的条件。我们通过从科学和技术信息办公室 (OSTI) 衍生的大约 100,000 科学和技术文档的语料库中派生标题,并使用组相对策略优化 (GRPO) 训练 Llama-3.1-8B-Instruct 来实例化该框架。通过基于 GRPO 的训练,,模型在保留的评分标准评估中实现了 $71.7\%$ 标准化奖励。 GRPO 调整的策略还在四个并非源自训练语料库的推理基准(GSM8K, MATH, GPQA Main, 和 GPQA Diamond)上改进了基本模型。这些结果证明,结构化, 基于文档的奖励可以提高坚持的评分标准性能,并诱导超出用于构建训练环境的语料库的可转移推理行为。

We argue that decomposing reward into weighted, verifiable criteria and using an LLM judge to score them provides a partial-credit optimization signal: instead of a binary outcome or a single holistic score, each response is graded along multiple task-specific criteria. We formalize \emph{rubric-grounded reinforcement learning (RL)}: a framework in which the policy is optimized against a structured, multi-criterion reward produced by a frozen LLM judge that conditions on auxiliary grounding the policy never sees. We instantiate the framework by deriving rubrics from an Office of Scientific and Technical Information (OSTI)-derived corpus of roughly 100,000 scientific and technical documents and training Llama-3.1-8B-Instruct with Group Relative Policy Optimization (GRPO). With GRPO-based training, the model achieves $71.7\%$ normalized reward on held-out rubric evaluation. The GRPO-tuned policy also improves over the base model on four reasoning benchmarks not derived from the training corpus -- GSM8K, MATH, GPQA Main, and GPQA Diamond. These results provide evidence that structured, document-grounded rewards can improve held-out rubric performance and induce transferable reasoning behaviors beyond the corpus used to construct the training environment.

科目:人工智能(cs.AI)

Subjects: Artificial Intelligence (cs.AI)