临床决策支持系统 (CDSS) 需要可仔细的, 可审计管道,以实现严格的, 可重复验证。然而,目前基于法学硕士的 CDSS 仍然基本上不透明。大多数 "open" 模型仅是开放式的, 发布参数,同时保留数据来源, 管理程序, 和确定模型行为的生成管道。完全开放的 (FO) 模型, 暴露了端到端的完整训练堆栈, 目前在医学中不存在。我们推出完全开放的 Meditron,,这是用于构建 LLM-CDSS, 的第一个完全开放的管道,其中包括经过临床医生审核的培训语料库,、可重复的数据构建和培训框架, 以及符合使用的评估协议。该语料库将八个公共医学 QA 数据集统一为标准化对话格式,并通过三个临床医生审查的综合扩展: 考试式 QA, 基于指南的 QA 扩大了覆盖范围,这些扩展源自 46,469 临床实践指南, 和临床小插图。该管道强制对教师一代,进行全系统净化,金标重新采样,并由四名医生小组进行端到端验证。我们使用法学硕士作为法官协议对专家编写的临床插图, 进行评估,并根据 204 名人类评估者进行校准。我们将该配方应用于五个 FO 基本模型 (Apertus-70B/8B-Instruct, OLMo-2-32B-SFT, EuroLLM-22B/9B-Instruct)。所有 MeditronFO 变体均优于其底座。 Apertus-70B-MeditronFO 在建立新的 FO SoTA 的综合医疗基准, 上比其基础 (47.2% 提高了 +6.6 点至 53.8%)。在法学硕士法官比较中,Gemma-3-27B-MeditronFO 优于 MedGemma,比率为 58.6%,并且在 HealthBench 上的表现优于 MedGemma((58% vs 55.9%))。这些结果表明,完全开放的管道可以在不牺牲可审核性或可重复性的情况下实现最先进的特定领域性能。
Clinical decision support systems (CDSS) require scrutable, auditable pipelines that enable rigorous, reproducible validation. Yet current LLM-based CDSS remain largely opaque. Most "open" models are open-weight only, releasing parameters while withholding the data provenance, curation procedures, and generation pipelines that determine model behavior. Fully Open (FO) models, which expose the complete training stack end-to-end, do not currently exist in medicine. We introduce Fully Open Meditron, the first fully open pipeline for building LLM-CDSS, comprising a clinician-audited training corpus, a reproducible data construction and training framework, and a use-aligned evaluation protocol. The corpus unifies eight public medical QA datasets into a normalized conversational format and expands coverage with three clinician-vetted synthetic extensions: exam-style QA, guideline-grounded QA derived from 46,469 clinical practice guidelines, and clinical vignettes. The pipeline enforces system-wide decontamination, gold-label resampling of teacher generations, and end-to-end validation by a four-physician panel. We evaluate using an LLM-as-a-judge protocol over expert-written clinical vignettes, calibrated against 204 human raters. We apply the recipe to five FO base models (Apertus-70B/8B-Instruct, OLMo-2-32B-SFT, EuroLLM-22B/9B-Instruct). All MeditronFO variants are preferred over their bases. Apertus-70B-MeditronFO improves +6.6 points over its base (47.2% to 53.8%) on aggregate medical benchmarks, establishing a new FO SoTA. Gemma-3-27B-MeditronFO is preferred over MedGemma in 58.6% of LLM-as-a-judge comparisons and outperforms it on HealthBench (58% vs 55.9%). These results show that fully open pipelines can achieve state-of-the-art domain-specific performance without sacrificing auditability or reproducibility.
科目: 人工智能 (cs.AI); 计算和语言 (cs.CL)
Subjects: Artificial Intelligence (cs.AI); Computation and Language (cs.CL)