电子健康记录(EHR)特征工程是临床研究的主要瓶颈,而AI,占数据科学家'工作量的39-45%。这在心力衰竭, 中尤其明显,心力衰竭影响了大约 670 万美国成年人,需要将零碎的 EHR 数据与特定疾病的, 基于指南的临床推理相结合。现有的基于规则和基于大型语言模型 (LLM) 的方法仅提供部分自动化,可维护性和证据可追溯性有限。我们开发了 Nimblemind 多代理系统 (nMAS), ,这是一个基于证据的 , 规则基础管道,用于自动心力衰竭特征工程 , ,并根据来自 9 个 EHR 源表的 500 条虚拟患者记录对其进行了评估。 nMAS 生成了 132 个结构化和 70 个评分聚合特征,,验证了结构完整性, 评分标准, 和出处,,并由受限制的法学硕士进行了审核。添加聚合特征后,HFrEF 的保留 AUROC 从 0.895 提高到 0.963,HFpEF 表型, 的保持 AUROC 从 0.870 提高到 0.910,并且基于 LLM 的独立证据支持和方法学健全性评估对这些特征进行了最高分 81.5% 的评分。这些结果证明了针对复杂心血管 EHR 数据, 进行自动化, 可审核特征工程的可行性,尽管评估仅限于单个机构队列并且需要外部验证。

Electronic health record (EHR) feature engineering is a major bottleneck in clinical research and AI, accounting for 39-45% of data scientists workload. This is especially pronounced in heart failure, which affects an estimated 6.7 million U.S. adults and requires integrating fragmented EHR data with disease-specific, guideline-based clinical reasoning. Existing rule-based and large language model (LLM)-based approaches offer only partial automation with limited maintainability and evidence traceability. We developed the Nimblemind Multi-Agent System (nMAS), an evidence-linked, rubric-grounded pipeline for automated heart-failure feature engineering, and evaluated it on 500 dummy patient records from nine EHR source tables. nMAS generated 132 structured and 70 rubric-scored aggregated features, verified for structural integrity, rubric compliance, and provenance, and audited by a restricted LLM. Adding the aggregated features improved held-out AUROC from 0.895 to 0.963 for HFrEF and 0.870 to 0.910 for HFpEF phenotyping, and an independent LLM-based rubric assessment of evidence support and methodological soundness scored the features at 81.5% of maximum points. These results demonstrate the feasibility of automated, auditable feature engineering for complex cardiovascular EHR data, though evaluation was limited to a single-institution cohort and external validation is needed.

科目: 人工智能 (cs.AI); 机器学习 (cs.LG)

Subjects: Artificial Intelligence (cs.AI); Machine Learning (cs.LG)