随着病原体基因组监测规模扩大,,瓶颈正在从数据生成转向分析。我们提出 BioSecBench-Surveillance, 是一个包含 100 项评估的可验证基准,测试 AI 代理是否可以从原始测序数据和监控上下文中推断出正确的分析流程。每次评估仅向代理提供人类分析师拥有的数据和上下文,,然后确定性地对其结构化答案进行评分。这些任务涵盖七个类别,,从分类学分类到基因工程检测,,涉及不同的样本类型和测序技术。在 16 个模型线束对 , 的 3,962 次可分级尝试中,最强的配置仅通过了大约一半。 Opus 4.8 的 PI 处于领先地位,为 50.2%,,在 83 项评估中,95% 置信区间为 40.1%,,95% 置信区间为 40.1%,,与 GPT-5.5 持平,Codex 为 50.2%,,95% 置信区间为 40.8%,,95% 置信区间为 40.8%,,其次是 Opus 4.7,PI 为 49.6%,, 95% 置信区间为 40.0% 到 59.2%,,Sonnet 4.6 的 PI 为 48.6%,,95% 置信区间为 38.9% 到 58.3%。即使代理调用了正确的工作流程,,他们的错误也来自于他们周围的选择,,例如引用,阈值,过滤器,以及要应用的标准化。 BioSecBench-Surveillance 提供了一个标准,用于衡量在下一次疫情爆发时是否可以信任代理执行基因组监测。

As pathogen genomic surveillance scales, the bottleneck is shifting from data generation to analysis. We present BioSecBench-Surveillance, a verifiable benchmark of 100 evaluations testing whether AI agents can infer the right analysis pipeline from raw sequencing data and surveillance context. Each evaluation gives an agent only the data and context a human analyst would have, then grades its structured answer deterministically. The tasks span seven categories, from taxonomic classification to genetic-engineering detection, across diverse sample types and sequencing technologies. Across 3,962 gradable attempts from sixteen model-harness pairs, the strongest configuration cleared only about half. Opus 4.8 with PI led at 50.2 percent, with a 95 percent confidence interval of 40.1 to 60.3 percent across 83 evaluations, tied with GPT-5.5 with Codex at 50.2 percent, with a 95 percent confidence interval of 40.8 to 59.6 percent, followed by Opus 4.7 with PI at 49.6 percent, with a 95 percent confidence interval of 40.0 to 59.2 percent, and Sonnet 4.6 with PI at 48.6 percent, with a 95 percent confidence interval of 38.9 to 58.3 percent. Even when agents invoked the correct workflows, their mistakes came from the choices around them, such as which references, thresholds, filters, and normalization to apply. BioSecBench-Surveillance provides a standard for measuring whether agents can be trusted to perform genomic surveillance when the next outbreak arrives.

科目:人工智能(cs.AI)

Subjects: Artificial Intelligence (cs.AI)