大型语言模型的最新进展导致各种任务, 的显着改进,包括数学推理,,用于评估模型 在逻辑推理和解决问题方面的智能。通过验证最终答案与真实答案的正确性,在数学推理基准上评估模型。此验证的常见方法是基于符号数学比较,,它无法泛化不同的数学表示和解决方案格式。在这项工作,中,我们为基于规则的符号数学比较提供了一个强大而灵活的替代方案。我们提出了一个基于 LLM 的评估框架,用于评估模型生成的答案,,从而能够跨不同的数学表示和答案格式进行准确评估。我们展示了两个流行框架, Lighteval 和 SimpleRL, 中符号求值的失败案例,并将它们与我们的方法, 进行比较,展示了对常用方法的明显改进。我们的框架能够实现更可靠的评估和基准测试,,从而实现更准确的性能监控,,这对于推进数学问题解决和智能系统非常重要。

Recent advancements in large language models have led to significant improvements across various tasks, including mathematical reasoning, which is used to assess models intelligence in logical reasoning and problem-solving. Models are evaluated on mathematical reasoning benchmarks by verifying the correctness of the final answer against a ground truth answer. A common approach for this verification is based on symbolic mathematics comparison, which fails to generalize across diverse mathematical representations and solution formats. In this work, we offer a robust and flexible alternative to rule-based symbolic mathematics comparison. We propose an LLM-based evaluation framework for evaluating model-generated answers, enabling accurate evaluation across diverse mathematical representations and answer formats. We present failure cases of symbolic evaluation in two popular frameworks, Lighteval and SimpleRL, and compare them to our approach, demonstrating clear improvements over commonly used methods. Our framework enables more reliable evaluation and benchmarking, leading to more accurate performance monitoring, which is important for advancing mathematical problem-solving and intelligent systems.

科目:人工智能(cs.AI)

Subjects: Artificial Intelligence (cs.AI)