公共人工智能评估通常被解读为终端排行榜,,但潜在的证据是由报告规则,基准修订,和缺失形成的选择性时间序列。 LiveBench 和 Open LLM Leaderboard v2 的重复公共档案作为主要纵向记录; LMArena 提供偏好压力测试; 且 GAIA 和 tau-bench 贡献有限的代理飞行员。 , 这些档案一起实例化了贝叶斯推理问题: 在固定的报告约定下, 一个构建的仅终端示例超过 $1{,}000$ 系统与两个终端前历史兼容, 产生 $23.03$ 或 $75.13$ 的时间,以达到上限的 $0.05$ 以内相同的终端尾部模型。在综合后验比较中,, 面向行动的诊断因观察方案而异。候选选择感知前沿模型未能实现综合恢复,客观档案预测,偏好转移,和不确定性校准;相应,固定审计门拒绝其更强的主张。存档和裁决协议重建公共评估历史, 隔离经过验证的时间边界, 并伪造不受支持的前沿主张。
Public AI evaluations are often read as terminal leaderboards, yet the underlying evidence is a selective time series shaped by reporting rules, benchmark revisions, and missingness. Repeated public archives for LiveBench and Open LLM Leaderboard v2 serve as the primary longitudinal record; LMArena provides a preference stress test; and GAIA and tau-bench contribute limited agentic pilots. Together, these archives instantiate a Bayesian inference problem: under a fixed reporting convention, one constructed terminal-only example over $1{,}000$ systems is compatible with two pre-terminal histories, yielding times of $23.03$ or $75.13$ to reach within $0.05$ of the ceiling under the same terminal-tail model. In synthetic posterior comparisons, action-facing diagnostics differ across observation regimes. The candidate selection-aware frontier model fails synthetic recovery, objective-archive prediction, preference transfer, and uncertainty calibration; correspondingly, fixed audit gates reject its stronger claims. An archive-and-adjudication protocol reconstructs public evaluation histories, isolates a verified timing boundary, and falsifies unsupported frontier claims.
科目: 人工智能 (cs.AI); 方法论 (stat.ME)
Subjects: Artificial Intelligence (cs.AI); Methodology (stat.ME)