大多数医疗人工智能基准测试都会衡量模型是否知道正确答案。 MedFailBench 提出了一个不同的问题: 哪个安全边界失败了? 我们提出了由临床医生构建的综合基准和失败图谱。该资源按严重程度从 1 到 5 以及安全门类型: 错过紧急升级, 不安全远程给药, 不安全出院保证, 证据伪造, 不安全协议执行, 和来源支持差距来标记医疗人工智能错误。当前公开版本 (v0.2.1) 包含由临床医生,审查的 44 个综合病例,其中包含严重程度注释,、公共拥抱面部空间源,、安全门分类法,、临床严重程度量规, 以及用于存档模型反应筛选运行的自动化管道。四十个案例具有已填充的安全门字段,,四个案例需要门完成。不包括患者数据, 临床验证声明, 或模型排名。 MedFailBench 在 Apache-2.0 和 CC-BY-4.0 下发布,并带有 Zenodo DOI https://doi.org/10.5281/zenodo.21205535。
Most medical AI benchmarks measure whether a model knows the correct answer. MedFailBench asks a different question: which safety boundary failed? We present a synthetic benchmark and failure atlas built by a clinician. The resource labels medical AI errors by severity from 1 to 5 and safety gate type: missed urgent escalation, unsafe remote dosing, unsafe discharge reassurance, evidence fabrication, unsafe protocol execution, and source support gap. The current public release (v0.2.1) contains 44 synthetic cases reviewed by a clinician, with severity annotations, a public Hugging Face Space source, a safety gate taxonomy, a clinical severity rubric, and an automated pipeline for archiving model response screening runs. Forty cases have a populated safety gate field, and four require gate completion. No patient data, clinical validation claims, or model rankings are included. MedFailBench is released under Apache-2.0 and CC-BY-4.0 and carries the Zenodo DOI https://doi.org/10.5281/zenodo.21205535.
科目: 人工智能 (cs.AI); 计算和语言 (cs.CL)
Subjects: Artificial Intelligence (cs.AI); Computation and Language (cs.CL)