多模型 LLM 系统(例如路由, 投票, 级联, 融合, 和混合代理)用于击败单模型准确性。我们表明,他们的收益受到该领域很少报告的数量的限制。对于输出为一个成员模型答案, 的任何策略,准确性不能超过 1 减去 beta,,其中 beta 是每个模型在同一查询上出错的比率。相反,, 通常的诊断, 平均成对误差相关性 rho, 无法识别具有相同边际的 beta: 误差规律,并且成对相关性可能具有不同的全错率。 Beta 上的 Clopper-Pearson 界给出了关于任何路由器, vote, 或级联在训练路由器之前可以提供的最大增益的有限样本证书。在来自 21 个提供商的 67 个模型, 中,四方校准的单因素模型仍然低估了开放式数学上完全错误的尾部:, 观察到的 beta 值为 0.052,而在完整的 67 模型高斯联结, 下观察到的 beta 值为 0.023,大约低估了, 的 2.5 倍,90% CI 为 1.7 至 3.4,k 等于 17。效果在执行分级代码, 上重复,其中 beta 为 0.079。以自由回答而不是多项选择的形式重新提出相同的 GPQA-Diamond 问题,会重新打开 tail, 的 beta 0.127 和由五位法官组成的小组,kappa 0.73 至 0.92, 在答案格式而不是主题中定位共同失败。在匹配质量, 低 rho 异构集成中击败了高 rho Self-MoA,,但在我们池中的可检查任务, 组合模型很少击败没有强大查询级路由信号的单个最佳模型。收益来自于模型在不同问题上失败,,而不是来自添加更多模型。

Multi-model LLM systems such as routing, voting, cascades, fusion, and mixture-of-agents are used to beat single-model accuracy. We show that their gain is capped by a quantity the field rarely reports. For any policy whose output is one member model answer, accuracy cannot exceed one minus beta, where beta is the rate at which every model is wrong on the same query. In contrast, the usual diagnostic, average pairwise error correlation rho, cannot identify beta: error laws with identical marginals and pairwise correlations can have different all-wrong rates. A Clopper-Pearson bound on beta gives a finite-sample certificate on the largest gain any router, vote, or cascade could deliver before training a router. Across 67 models from 21 providers, a tetrachoric-calibrated single-factor model still underprices the all-wrong tail: on open-ended mathematics, observed beta is 0.052 versus 0.023 under the full 67-model Gaussian copula, about 2.5 times underpricing, with 90 percent CI 1.7 to 3.4 and k equals 17. The effect recurs on execution-graded code, where beta is 0.079. Re-asking the same GPQA-Diamond questions in free-response rather than multiple-choice form reopens the tail, with beta 0.127 and a five-judge panel with kappa 0.73 to 0.92, locating co-failure in answer format rather than subject. At matched quality, low-rho heterogeneous ensembles beat high-rho Self-MoA, but on checkable tasks in our pool, combining models rarely beats the single best model without a strong query-level routing signal. Gains come from models failing on different questions, not from adding more models.

科目: 人工智能 (cs.AI); 机器学习 (cs.LG)

Subjects: Artificial Intelligence (cs.AI); Machine Learning (cs.LG)