肝细胞癌 (HCC) 是一种常见的恶性肿瘤,也是癌症相关死亡的主要原因。当前的指南和分期系统提供了粗略的类别,,但经常错过阶段内的异质性和电子病历中的临床背景(EMRs)。我们提出 HCC-STAR (肝细胞癌分期, 治疗和预后), 是一种临床对齐的大型语言模型,可读取常规 EMR 叙述并联合输出基于风险评分的分期, 排名一致的指南治疗方法,并具有基于证据的基本原理, 和个体化生存估计。我们从 SEER 中精选了约 30,000 个 HCC 病例,并使用经临床医生验证的, 基于提示的增强工作流程将其扩展为 EMR 式叙述训练数据。在这个语料库,上,我们开发了一个知识一致的推理框架,该框架通过可逐步验证的复合奖励,进行了优化,超越了临床指南的文本级记忆。在来自中国 12 家医院的 6,668 名患者的多中心队列中,与临床指南和包括 GPT-5 和 Gemini-2.5 Pro 在内的竞争模型, 相比,, HCC-STAR 在治疗推荐和风险分层方面取得了最先进的表现。假设的总体生存分析显示,遵守 HCC-STAR 建议, 的中位生存期为 51 个月,而 BCLC 和 CNLC 的中位生存期为 29 和 32 个月。在以临床医生为中心的评估中,, 盲法肝胆专家将 HCC-STAR的推理和基于证据的理由评为值得信赖。该模型在治疗准确性方面超越了住院医生和主治医生,并在作为助手时帮助医生更快地做出更准确的决策。这些发现支持 HCC-STAR 作为 HCC 风险分层和精准治疗的可靠且可验证的决策支持系统。
Hepatocellular carcinoma (HCC) is a common malignancy and a leading cause of cancer-related mortality. Current guidelines and staging systems provide coarse categories, but often miss within-stage heterogeneity and the clinical context in electronic medical records (EMRs). We present HCC-STAR (Hepatocellular Carcinoma Staging, Treatment And pRognosis), a clinically aligned large language model that reads routine EMR narratives and jointly outputs risk score-based staging, ranked guideline-consistent treatments with evidence-based rationales, and individualized survival estimates. We curated about 30,000 HCC cases from SEER and expanded them into EMR-style narrative training data using a clinician-validated, prompt-based augmentation workflow. On this corpus, we developed a knowledge-aligned reasoning framework optimized with a step-verifiable composite reward, moving beyond text-level memorization of clinical guidelines. In a multi-center cohort of 6,668 patients from 12 hospitals in China, HCC-STAR achieved state-of-the-art performance in treatment recommendation and risk stratification compared with clinical guidelines and competitive models, including GPT-5 and Gemini-2.5 Pro. Hypothetical overall-survival analysis showed a median survival of 51 months under adherence to HCC-STAR recommendations, compared with 29 and 32 months under BCLC and CNLC. In clinician-centric evaluations, blinded hepatobiliary specialists rated HCC-STAR的 reasoning and evidence-based justifications as trustworthy. The model surpassed resident and attending physicians in treatment accuracy and helped physicians make more accurate decisions faster when used as an assistant. These findings support HCC-STAR as a reliable and verifiable decision-support system for risk stratification and precision therapy in HCC.
科目:人工智能(cs.AI)
Subjects: Artificial Intelligence (cs.AI)