客观的。临床人工智能文档系统需要临床有效,、经济可行, 且对迭代变化敏感的评估方法。每个评分实例都需要专家评审的方法对于安全, 迭代部署来说太慢且昂贵。我们提出了一种针对特定病例的,临床医生编写的用于临床人工智能评估的评估标准方法,并检查法学硕士生成的评估标准是否可以近似临床医生的一致性。材料和方法。 20 名临床医生为 823 个临床病例 (736 个真实世界, 87 个合成) 编写了 1,646 评分标准,涵盖初级保健, 精神病学, 肿瘤学, 和行为健康。通过确认基于法学硕士的评分代理对临床医生首选输出的评分始终高于被拒绝的输出,对每个评分标准进行了验证。针对所有案例对临床医生使用的七个版本的 EHR 嵌入式人工智能代理进行了评估。结果。临床医生编写的评分标准可有效地区分高质量和低质量的输出(中位得分差距: 82.9%),并且得分稳定性高(中位范围: 0.00%)。中位分数从 84% 提高到 95%。在后来的实验中,, 临床医生-LLM 排名一致性 (tau: 0.42-0.46) 匹配或超过了临床医生-临床医生一致性 (tau: 0.38-0.43),,这归因于上限压缩和 LLM 评分标准的改进。讨论。这种融合支持将法学硕士的标准与临床医生编写的标准相结合。以大约 1,000 倍的较低成本, LLM 规则可实现更大的评估覆盖率,,同时持续的临床作者身份根据专家判断进行评估。上限压缩对未来评估者间一致性研究提出了方法论挑战。结论。针对具体案例的评估标准为临床人工智能评估提供了一条途径,可以保留专家判断,同时以降低三个数量级的成本实现自动化。临床医生撰写的评估标准建立了验证法学硕士评估标准的基线。

Objective. Clinical AI documentation systems require evaluation methodologies that are clinically valid, economically viable, and sensitive to iterative changes. Methods requiring expert review per scoring instance are too slow and expensive for safe, iterative deployment. We present a case-specific, clinician-authored rubric methodology for clinical AI evaluation and examine whether LLM-generated rubrics can approximate clinician agreement. Materials and Methods. Twenty clinicians authored 1,646 rubrics for 823 clinical cases (736 real-world, 87 synthetic) across primary care, psychiatry, oncology, and behavioral health. Each rubric was validated by confirming that an LLM-based scoring agent consistently scored clinician-preferred outputs higher than rejected ones. Seven versions of an EHR-embedded AI agent for clinicians were evaluated across all cases. Results. Clinician-authored rubrics discriminated effectively between high- and low-quality outputs (median score gap: 82.9%) with high scoring stability (median range: 0.00%). Median scores improved from 84% to 95%. In later experiments, clinician-LLM ranking agreement (tau: 0.42-0.46) matched or exceeded clinician-clinician agreement (tau: 0.38-0.43), attributable to both ceiling compression and LLM rubric improvement. Discussion. This convergence supports incorporating LLM rubrics alongside clinician-authored ones. At roughly 1,000 times lower cost, LLM rubrics enable substantially greater evaluation coverage, while continued clinical authorship grounds evaluation in expert judgment. Ceiling compression poses a methodological challenge for future inter-rater agreement studies. Conclusion. Case-specific rubrics offer a path for clinical AI evaluation that preserves expert judgment while enabling automation at three orders lower cost. Clinician-authored rubrics establish the baseline against which LLM rubrics are validated.

科目: 人工智能 (cs.AI); 计算和语言 (cs.CL)

Subjects: Artificial Intelligence (cs.AI); Computation and Language (cs.CL)