临床实践不是从列举的选项中选择答案:,医生逐渐收集异质信息,并在不确定的情况下做出连续的, 不可逆转的决定。静态基准无法探测,现有的交互式医疗基准至少会在其中一个基准上做出妥协。我们提出了 ClinEnv, 一个交互式基准,该基准在我们称之为纵向住院模拟的范式下评估法学硕士作为主治医生与真实住院患者的入院情况。每个案例都会自动构建为决策阶段;的有序序列,在每个阶段,模型必须主动查询四个专门代理,然后再进行药物,程序,和诊断。 ClinEnv 对模型通过确定性本体匹配, 决定的内容, 以及它如何收集信息进行评分。在 7 个模型, 中,最强的模型仅达到 0.31 决策 F1,,并且结果质量与过程质量急剧脱钩。困难集中在管理决策和后期,,其中模型恢复出院诊断的可靠性远高于管理行动(0.51 vs. 0.17 F1),并随着病例的进展继续发出冗余查询。 ClinEnv 使这种信息获取差距, 对于仅结果评估, 来说是不可见的,可直接测量。

Clinical practice is not the selection of an answer from enumerated options: a physician gathers heterogeneous information incrementally and commits to sequential, irreversible decisions under uncertainty. Static benchmarks cannot probe and existing interactive medical benchmarks each compromise on at least one of them. We present ClinEnv, an interactive benchmark that evaluates LLMs as attending physicians over real inpatient admissions under a paradigm we term Longitudinal Inpatient Simulation. Each case is automatically constructed into an ordered sequence of decision stages; at every stage the model must actively query four specialized agents before committing to medications, procedures, and diagnoses. ClinEnv scores both what the model decides, through deterministic ontology-grounded matching, and how it gathers information. Across seven models, the strongest reaches only 0.31 decision F1, and outcome quality is sharply decoupled from process quality. Difficulty concentrates in management decisions and later stages, where models recover discharge diagnoses far more reliably than management actions (0.51 vs. 0.17 F1) and continue to issue redundant queries as cases progress. ClinEnv makes this information-acquisition gap, invisible to outcome-only evaluation, directly measurable.

科目: 人工智能 (cs.AI); 计算和语言 (cs.CL); 新兴技术 (cs.ET); 多代理系统 (cs.MA)

Subjects: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Emerging Technologies (cs.ET); Multiagent Systems (cs.MA)