LLM 的知识基准面临三个问题: 扩展驱动的设计,没有实现学科代表性; 固定支付注释,允许懒惰的共识; 和有限测试预算下未经审计的排名不稳定。我们引入了 KINA,,这是一个涵盖 261 个细粒度学科的 899 项基准,,有两个正式结果。首先,,我们将代表性作为专家引出的锚的覆盖式目标,并通过代理, 操作学科代表性,产生 (1-1/e) 贪婪近似(命题 1); 该保证适用于代理, 而不是人口代表性。第二, 我们证明了奖金酒吧锦标赛在发布评论质量, 中弱于 FOSD 主导固定支付,激励兼容性阈值 B > Delta C / Delta p_min ( 定理 1)。通过评估 13 个实验室的 42 个模型,,顶级模型, Gemini-3.1-Pro-Preview, 达到 53.17%,,其次是 Claude-Opus-4.6,达到 49.92%,GPT-5.4 达到 48.55%,,在饱和度以下还有很大的空间。完整的排行榜显示了分层结构,而不是平滑的总订单:,一个小前沿层位于 48%, 之上,一个密集的强模型层跨越大约 38-45%,,而低性能模型仅略高于 10% 机会基线。在五种工具使用评估, 中,工具增强加起来高达 5.17 分,不同模型的增益差异很大。我们报告引导排名稳定性统计数据,以使有限预算方差明确并阻止对相邻排名的过度解释。
Knowledge benchmarks for LLMs face three issues: scaling-driven designs that do not operationalize disciplinary representativeness; flat-payment annotation that permits lazy consensus; and unaudited ranking instability under bounded test budgets. We introduce KINA, an 899-item benchmark across 261 fine-grained disciplines, with two formal results. First, we cast representativeness as a coverage-style objective over expert-elicited anchors and operationalize disciplinary representativeness through a proxy, yielding a (1-1/e) greedy approximation (Proposition 1); the guarantee applies to the proxy, not to population representativeness. Second, we prove a bonus-on-bar tournament weakly FOSD-dominates flat payment in released-review quality, with incentive-compatibility threshold B > Delta C / Delta p_min (Theorem 1). Evaluating 42 models from 13 labs, the top model, Gemini-3.1-Pro-Preview, reaches 53.17%, followed by Claude-Opus-4.6 at 49.92% and GPT-5.4 at 48.55%, leaving substantial headroom below saturation. The full leaderboard shows a tiered structure rather than a smooth total order: a small frontier tier lies above 48%, a dense strong-model tier spans roughly 38-45%, and low-performing models remain only modestly above the 10% chance baseline. Tool augmentation adds up to 5.17 points across the five tool-use evaluations, with gains varying substantially across models. We report bootstrap ranking-stability statistics to make bounded-budget variance explicit and to discourage over-interpretation of adjacent ranks.
科目:人工智能(cs.AI)
Subjects: Artificial Intelligence (cs.AI)