大型语言模型的标准评估假设在推理条件下模型排名稳定。我们通过改变令牌生成预算,来挑战这一假设,即,模型可以在七个级别(64--4,096),上产生,的最大令牌,在三个推理基准(56,476推论)上评估四个模型。我们报告了四个发现: (i) 3--19% 的项目表现出非单调行为(即使在控制截断,之后,准确性也会随着更多预算),而降低,并且这种现象是特定于模型的(跨模型重叠: 6--14%)。 (ii) 在所有基准上,模型排名在预算中逆转 ($p {<} 0.01$, McNemar)。 (iii) Oracle 分析显示,模型互补性高达 $+27.8$pp,,在预算有限的情况下最为明显。 (iv) 预算感知路由器捕获 14.1% 的 Oracle 差距跨域; 预算功能有助于域内 ($+1.6$ 到 $+5.7$pp),但特定于域并损害传输 ($-1.2$pp)。这些结果支持以预算为条件的评估方案。

Standard evaluation of large language models assumes stable model rankings across inference conditions. We challenge this assumption by varying the token generation budget, i.e., the maximum tokens a model may produce, across seven levels (64--4,096), evaluating four models on three reasoning benchmarks (56,476 inferences). We report four findings: (i) 3--19% of items exhibit non-monotone behavior (accuracy decreasing with more budget), even after controlling for truncation, and this phenomenon is model-specific (cross-model overlap: 6--14%). (ii) Model rankings reverse across budgets on all benchmarks ($p {<} 0.01$, McNemar). (iii) Oracle analysis reveals model complementarity up to $+27.8$pp, most pronounced at constrained budgets. (iv) A budget-aware router captures 14.1% of the oracle gap cross-domain; budget features help within-domain ($+1.6$ to $+5.7$pp) but are domain-specific and hurt transfer ($-1.2$pp). These results argue for budget-conditioned evaluation protocols.

科目: 人工智能 (cs.AI); 计算和语言 (cs.CL)

Subjects: Artificial Intelligence (cs.AI); Computation and Language (cs.CL)