由大型语言模型 (LLMs) 提供支持的面向消费者的健康聊天机器人越来越多地用于症状评估。然而, 聊天机器人的开发和评估通常依赖于合作的, 清晰的, 模拟患者。我们分析了 2,053 真实的患者与聊天机器人对话,发现不同用户的沟通模式和情绪表达差异很大。我们开发了一个病人模拟器,可以分别模拟临床内容,情绪状态,对话策略,和沟通风格。在由 15 名人类评分者, 进行的受图灵启发的现实主义评估中,模拟对话与真实对话几乎无法区分,,人类评分者的准确率达到 55%。我们在 1,164 个临床医生分级病例, 中使用了 5 个不同的患者角色, 来评估四位法学硕士在紧急情况评估中的表现。我们发现沟通方式可以显着改变分类结果。以患者为中心的对话人工智能必须适应通信多样性:系统,该系统是为理想化的,而不是现实的,交互而设计的,当部署在现实世界中时,系统可能会表现不佳并扩大健康差距。

Consumer-facing health chatbots powered by large language models (LLMs) are increasingly used for symptom assessment. However, chatbot development and evaluation often rely on cooperative, articulate, simulated patients. We analysed 2,053 real patient-chatbot conversations and found that communication patterns and expression of emotions vary widely across users. We developed a patient simulator that separately models clinical content, emotional state, conversational strategy, and communication style. In a Turing-inspired evaluation of realism with 15 human graders, simulated conversations were nearly indistinguishable from real ones, with human graders achieving an accuracy of 55%. We used five distinct patient personae, across 1,164 clinician-graded cases, to evaluate the performance of four LLMs in urgency assessment. We found that communication style can significantly alter triage outcomes. Patient-centred conversational artificial intelligence must accommodate communication diversity: systems designed for idealised, rather than realistic, interactions risk underperforming and amplifying health disparities when deployed in the real world.

科目: 人工智能 (cs.AI); 计算和语言 (cs.CL)

Subjects: Artificial Intelligence (cs.AI); Computation and Language (cs.CL)