企业人工智能代理的综合临床基准可以通过现有的实用程序检查,但在结构上仍然不切实际,,特别是在难以访问操作数据的隐私敏感的医疗保健环境中。我们研究如何在不破坏实践中已使用的下游实用程序检查的情况下改进此类基准。我们将基准修订制定为受实用程序限制的现实性改进: 数据集更改应增加现实性,同时保持在可操作实用程序地板之上。我们在护理差距基准上实例化了这一想法,该基准源自 Synthea 生成的患者,通过演示电子健康记录工作流程进行锻炼,然后由与操作数据相同的下游管道进行处理。真实性是通过缺失结构,简单性,结构合理性,和群体对齐来衡量的。基线基准非常薄: 采样对缺失率为 79.44%, 只有 12.75% 行可操作, 38.94% 的患者具有零可操作措施, 前三名令牌集中度达到 100.0%。两个确定性修订改进了这些面板,同时保持在当前实用底层,之上,而天真的致密化控制保留了不切实际的模板。我们进一步表明,内部基准现实性和对总体操作参考的源保真度是相关但不同的目标。这些结果表明,应明确优化综合基准质量,,并将效用视为一种约束,而不是作为现实主义的充分证据。

Synthetic clinical benchmarks for enterprise AI agents can pass existing utility checks and still remain structurally unrealistic, especially in privacy-sensitive healthcare settings where operational data are hard to access. We study how to improve such benchmarks without breaking the downstream utility checks already used in practice. We formulate benchmark revision as utility-constrained realism improvement: dataset changes should increase realism while staying above an operational utility floor. We instantiate this idea on a care-gap benchmark derived from Synthea-generated patients exercised through demonstration electronic health record workflows and then processed by the same downstream pipeline as operational data. Realism is measured through missingness structure, simplicity, structural plausibility, and population alignment. The baseline benchmark is extremely thin: sampled-pair missingness is 79.44%, only 12.75% of rows are actionable, 38.94% of patients have zero actionable measures, and top-three token concentration reaches 100.0%. Two deterministic revisions improve these panels while remaining above the current utility floor, whereas a naive densification control preserves unrealistic templating. We further show that internal benchmark realism and source fidelity to an aggregate operational reference are related but distinct objectives. These results suggest that synthetic benchmark quality should be optimized explicitly, with utility treated as one constraint rather than as sufficient evidence of realism.

科目: 人工智能 (cs.AI); 数据库 (cs.DB); 机器学习 (cs.LG)

Subjects: Artificial Intelligence (cs.AI); Databases (cs.DB); Machine Learning (cs.LG)