深入研究,,其中代理搜索开放网络, 收集证据, 并通过扩展推理得出答案, 是前沿语言模型的一个突出用例。前沿深度研究产品在现有基准,上得分很高,因此很难仅从当前的评估数据中区分出它们的能力。我们引入 DeepWeb-Bench, 一个深度研究基准,它比当前前沿的现有基准要难得多。困难来自于数据本身的三个属性:每项任务需要大量证据收集,跨源核对,和长期多步推导。我们将这三个困难来源表示为四个功能系列(检索,推导,推理,和校准),并报告按系列划分的结果。每个参考答案都附有源出处记录,具有四个披露级别和跨源检查(如果可用), 使分数更容易根据基础证据进行审核。我们在九个前沿模型上评估 DeepWeb-Bench,并报告三个结果: (1) 检索不是瓶颈,,因为检索失败仅占错误的 12-14%,而推导和校准失败则占 70% 以上25; (2) 强模型和弱模型以不同的方式失败, 强模型 错误以不完整推导和弱模型为主'幻觉精度; 和 (3) 模型表现出跨领域的真正专业化,,跨模型一致性仅为 rho = 0.61,每个案例的分歧达到 18.8 个百分点。公共基准测试版本包括数据, rubrics, 和评估代码。
Deep research, in which an agent searches the open web, collects evidence, and derives an answer through extended reasoning, is a prominent use case for frontier language models. Frontier deep research products score high on existing benchmarks, making it difficult to distinguish their capabilities from current evaluation data alone. We introduce DeepWeb-Bench, a deep research benchmark that is substantially harder than existing benchmarks for the current frontier. Difficulty comes from three properties of the data itself: each task requires massive evidence collection, cross-source reconciliation, and long-horizon multi-step derivation. We represent these three sources of difficulty as four capability families (Retrieval, Derivation, Reasoning, and Calibration) and report results sliced by family. Every reference answer is accompanied by a source-provenance record with four disclosure levels and cross-source checks where available, making scores easier to audit against the underlying evidence. We evaluate DeepWeb-Bench on nine frontier models and report three findings: (1) retrieval is not the bottleneck, as retrieval failures account for only 12-14% of errors while derivation and calibration failures account for over 70%; (2) strong and weak models fail in qualitatively different ways, with strong models errors dominated by incomplete derivation and weak models by hallucinated precision; and (3) models exhibit genuine specialization across domains, with cross-model agreement of only rho = 0.61 and per-case disagreement reaching 18.8 percentage points. The public benchmark release includes the data, rubrics, and evaluation code.
科目:人工智能(cs.AI)
Subjects: Artificial Intelligence (cs.AI)