AI 基准越来越多地利用项目级统计模型,,特别是项目响应理论(IRT), 来估计模型能力, 排名系统, 选择信息丰富的示例, 并诊断基准质量。然而,, AI 基准数据通常偏离人类测试的数据体系,,标准 IRT 估计工具最初是为其开发的: 基准通常涉及较少的评估模型, 更多的项目, 和可能倾斜的, 集群, 或多模式的能力分布。我们研究了这些机制不匹配如何挑战用于人工智能评估的 IRT 模型的可靠性。使用源自六个广泛使用的 LLM 基准, 的项目参数和能力分布,我们模拟了三种常见 IRT 模型下的响应矩阵,并比较了最近基准研究中使用的四种估计工具: 边际最大似然, 马尔可夫链蒙特卡罗, 变分推理, 和神经伪暹罗估计器。在 18,000 模拟条件, 中,我们系统地评估了计算可行性, 可扩展性, 以及关于模型排名, 预测性能, 和项目特征的 IRT 推断的可靠性。结果表明,经典估计器在大型基准设置, 中可能变得不可行,而可扩展估计器可能会使用小型或非正态分布模型集产生不可靠的项目级别和排名推断。这项研究确定了潜在特征模型何时可靠地支持人工智能基准测试声明,或存在扭曲人工智能基准声明,的风险,以及需要哪些样本量和诊断才能可靠使用。

AI benchmarks increasingly leverage item-level statistical models, particularly item response theory (IRT), to estimate model capabilities, rank systems, select informative examples, and diagnose benchmark quality. However, AI benchmark data often departs from the data regime of human testing, for which standard IRT estimation tools were originally developed: benchmarks typically involve fewer evaluated models, far more items, and capability distributions that may be skewed, clustered, or multimodal. We examine how these regime mismatches challenge the reliability of IRT modeling for AI evaluation. Using item parameters and capability distributions derived from six widely used LLM benchmarks, we simulate response matrices under three common IRT models and compare four estimation tools used in recent benchmark studies: marginal maximum likelihood, Markov chain Monte Carlo, variational inference, and a neural pseudo-Siamese estimator. Across 18,000 simulation conditions, we systematically evaluate computational feasibility, scalability, and the reliability of IRT inferences about model rankings, predicted performance, and item characteristics. Results show that classical estimators can become infeasible in large benchmark settings, whereas scalable estimators can produce unreliable item-level and ranking inferences with small or non-normally distributed model sets. This study identifies when latent trait models reliably support or risk distorting AI benchmarking claims, and what sample sizes and diagnostics are needed for trustworthy use.

科目:人工智能(cs.AI)

Subjects: Artificial Intelligence (cs.AI)