人工智能 (AI) 代理有望通过压缩解释和决策循环, 来加速药物发现,但实际部署需要对实际项目决策进行可信评估。我们推出了 TherapeuticsBench 临床前药理学 (TxBench-PP),,它是小分子临床前药理学的可验证基准,也是跨药物发现阶段和治疗方式的更广泛 TherapeuticsBench 工作的第一个重点部分。 TxBench-PP 测试代理是否可以从现实世界的分析数据而不是从文献中记忆的事实中恢复准确的结论。该基准包含按程序阶段,测定类型,和任务结构,索引的100项评估,涵盖作用机制(MoA)和药效学(PD)推理,化合物-靶点参与,因果靶点验证,可开发性和安全性,和转化功效。代理接收真实的工作流程快照, 在编码环境, 中检查文件并返回确定性分级的结构化答案。在 16 个模型线束配置, 中,包括 11 个模型和 4,800 轨迹,,没有系统能够可靠地恢复临床前药理学决策。最强配置, Claude Opus 4.8 / Pi, 通过了 59.3\% 的端点尝试 (178/300; 95\% CI, 51.1-67.6), 其次是 GPT-5.5 / Pi 在 55.3\% (166/300; 47.0-63.6)。

Artificial intelligence (AI) agents promise to accelerate drug discovery by compressing interpretation and decision-making loops, but practical deployment requires trusted evaluation on realistic program decisions. We introduce TherapeuticsBench Preclinical Pharmacology (TxBench-PP), a verifiable benchmark for small-molecule preclinical pharmacology and the first focused slice of a broader TherapeuticsBench effort across drug-discovery stages and therapeutic modalities. TxBench-PP tests whether agents can recover accurate conclusions from real-world assay data rather than memorized facts from literature. The benchmark contains 100 evaluations indexed by program stage, assay type, and task structure, spanning mechanism-of-action (MoA) and pharmacodynamic (PD) reasoning, compound-target engagement, causal target validation, developability and safety, and translational efficacy. Agents receive realistic workflow snapshots, inspect files in a coding environment, and return structured answers graded deterministically. Across 16 model-harness configurations, comprising 11 models and 4,800 trajectories, no system reliably recovered preclinical pharmacology decisions. The strongest configuration, Claude Opus 4.8 / Pi, passed 59.3\% of endpoint attempts (178/300; 95\% CI, 51.1-67.6), followed by GPT-5.5 / Pi at 55.3\% (166/300; 47.0-63.6).

科目: 人工智能 (cs.AI); 机器学习 (cs.LG)

Subjects: Artificial Intelligence (cs.AI); Machine Learning (cs.LG)