在设计协助用户进行疾病诊断的人工智能系统时,一刀切的方法可能不是的最佳策略。
A one-size-fits-all approach likely isn’t the best strategy when designing artificial intelligence systems that assist users in disease diagnosis.
麻省理工学院和其他地方的研究人员进行的一项新研究发现,人工智能辅助通常提高了非专家和临床医生诊断皮肤病的准确性,人工智能可解释性方法根据用户的知识水平产生不同的影响’。
A new study by researchers at MIT and elsewhere found that, while AI assistance generally improved the accuracy of non-experts and clinicians in diagnosing skin diseases, AI explainability methods had different impacts depending on the users knowledge level.
可解释的 AI 方法通过描述或验证模型的 的决策,帮助用户知道何时信任模型的 的预测。例如,,模型可能使用热图突出显示对其诊断最重要的图像区域,或使用大型语言模型(LLM) 以简单语言解释预测。
Explainable AI methods help users know when to trust a model的 predictions by describing or validating the model的 decision-making. For instance, a model might use a heat map to highlight image regions that were most important in its diagnosis or a large language model (LLM) to explain the prediction in plain language.
在这项研究中,, 研究人员在有或没有不同可解释的人工智能系统的帮助下测试了非专家和初级保健提供者在皮肤病诊断, 方面的情况。
In this study, researchers tested non-experts and primary care providers in skin disease diagnosis, with and without the help of different explainable AI systems.
他们发现非专家的诊断准确性提高了,,但这很大程度上是由于对人工智能系统的尊重。非专家相信基于法学硕士的解释,无论它们是对还是错,,并发现当解释含糊或笼统时更有说服力。
They found that non-experts diagnostic accuracy improved, but it was largely due to deference to the AI system. Non-experts trusted LLM-based explanations whether they were right or wrong, and found explanations more convincing when they were vague or generic.
By contrast, clinicians were not tripped up by incorrect AI assistance and performed best when given only a model的 prediction, with no accompanying explanation.
“These findings are important as patients increasingly turn to AI to help with their health care.我们的研究结果表明,当可解释的人工智能模型给出错误的输出时,那些医学知识最少的人最有可能误入歧途,”,斯坦福大学生物医学数据科学和皮肤病学的合著者兼助理教授 Roxana Daneshjou, 说。
“These findings are important as patients increasingly turn to AI to help with their health care. Our findings show that those with the least medical knowledge are most likely to be led astray when explainable AI models give an erroneous output,” says Roxana Daneshjou, a co-author and assistant professor of biomedical data science and dermatology at Stanford University.
These results underscore the importance of building AI systems with users in mind and of developing explainability methods that encourage critical thinking rather than overreliance on the model, the researchers say.
Ghassemi, Xu, 和 Daneshjou 与许多作者, 一起撰写了这篇论文,其中包括麻省理工学院研究生张浩然, 本科生 Reina Wang, 和 Luis Soenksen 博士 ’20,(Jameel Clinic, 的研究附属机构)以及临床医生和研究人员。对这项工作的描述今天发表在《自然医学》杂志上。
Ghassemi, Xu, and Daneshjou are joined on the paper by many authors, including MIT graduate student Haoran Zhang, undergraduate Reina Wang, and Luis Soenksen PhD ’20, a research affiliate at the Jameel Clinic, along with clinicians and researchers. A description of the work appears today in Nature Medicine.
多个 FDA 批准的人工智能接口被用来帮助临床医生识别医学图像中的皮肤状况,,作为简化早期诊断的一种方式。除了提供图像中是否存在疾病的预测, 之外,这些工具还经常使用解释模型 决策的几种方法之一。
Several FDA-approved AI interfaces are being used to help clinicians identify skin conditions in medical images, as a way to streamline early diagnosis. In addition to providing a prediction of whether disease is present in the image, these tools often use one of several methods that explain the model的 decision-making.
同时,非专家可以使用人工智能驱动的搜索引擎自行进行数字诊断,根据用户提示预测皮肤疾病。 These systems often use LLMs to explain the model的 prediction in simpler terms.
At the same time, non-experts can perform digital diagnosis on their own using AI-powered search engines that predict skin diseases based on user prompts. These systems often use LLMs to explain the model的 prediction in simpler terms.
研究人员探索了这些可解释的人工智能工具对初级保健医生和皮肤病检测非专家的影响和潜在好处。 They tested users by showing them medical images plus an AI prediction of skin disease, employing different explainable AI approaches.
The researchers explored the effects and potential benefits of these explainable AI tools on primary care physicians and non-experts in dermatological disease detection. They tested users by showing them medical images plus an AI prediction of skin disease, employing different explainable AI approaches.
这些方法包括: 人工智能预测和置信度,无需解释, 一种提供相似图像以强化其预测的方法, 一种基于热图的方法,突出显示重要图像区域, 以及一个法学硕士,用简单语言解释模型的 推理。
These approaches included: an AI prediction and confidence level with no explanation, a method that provides similar images to reinforce its prediction, a heat map-based approach that highlights important image regions, and an LLM that explains the model的 reasoning in plain language.
非专家的任务是在有或没有可解释人工智能的帮助下确定皮肤痣的图像是否癌变,。临床医生面临着更具挑战性的任务,即提供皮肤病的鉴别诊断。
Non-experts were tasked with deciding whether an image of a skin mole was cancerous, with and without the help of explainable AI. Clinicians were given the more challenging task of providing a differential diagnosis of dermatological disease.
The researchers found that all explainable AI approaches improved the accuracy of non-experts, mostly because the tools helped users diagnose non-cancerous moles.
此外, 当他们采用公平约束模型来消除对深色肤色的偏见, 时,系统显着提高了准确性并减少了基于肤色的诊断差异。
In addition, when they employed a fairness-constrained model designed to combat bias against darker skin tones, the system significantly improved accuracy and reduced diagnostic disparities based on skin tone.
“But the reason non-expert users are better is because they are more reliant on the models. When the model is wrong, it hurts performance more than it helps performance when the model is right. We were just able to train very good AI models for this setting,” Ghassemi says.
This deference effect is largest with LLM explanations, and users were more confident about their wrong answers when aided by an LLM.
On the other hand, clinicians were resilient to incorrect AI explanations and, of all the explainability methods, LLMs boost their accuracy the least.
“这实际上取决于每个小组如何使用解释。临床医生心里已经有了诊断,并根据自己的训练, 检查人工智能,这样就会发现错误的解释。同时,非专家可以首先使用完全相同的解释来形成意见,,因此看似合理的,听起来充满信心的理由可以将他们引向错误的答案。同样的工具最终成为一个用户的资产,而对另一个用户来说则是负债,” Xu 说。
“It really comes down to how each group uses the explanation. A clinician already has a diagnosis in mind and checks the AI against their own training, so a bad explanation gets caught. Meanwhile, a non-expert can use that exact same explanation to form an opinion in the first place, so a plausible, confident-sounding rationale can pull them toward the wrong answer. The same tool ends up being an asset for one user and a liability for another,” Xu says.
When the researchers dug deeper, they found that users who were most deferential to AI assistance were the worst performers on the task without the help of AI.
They also found that the time at which users were presented with AI explanations influenced their behavior. If an explanation is given first, before the user can perform the diagnosis on their own, they tend to become more deferential to the model.
In addition, AI systems outperformed humans when the presentation of disease was subtle, but humans performed much better if there are atypical symptoms or unrelated features in an image.
总而言之,这些结果表明,可解释的人工智能可能会导致对模型的过度依赖,并导致用户盲目遵循人工智能的建议,即使这些建议是错误的。
Taken together, these results indicate that explainable AI can cause overreliance on models and lead users to blindly follow AI recommendations even when they are wrong.
与其使用 LLM 生成更详细的解释,,不如强制用户首先给出诊断假设,,然后提供基于 AI 的建议来突出显示其他可能的条件以供考虑,这可能会更有效。
Rather than using LLMs to generate more detailed explanations, it might be more effective to force users to give a diagnostic hypothesis first, then provide an AI-based suggestion to highlight other possible conditions for consideration.
“We really want AI to improve creativity and either upskill or fill in gaps where users are missing subtle presentations. Otherwise, we risk engaging automation bias and then, when the model is wrong, users can’t recover,” Ghassemi says.
这项研究的,部分,由国家科学基金会,施密特科学,、国家经济研究局,和哥伦比亚大学资助。
This research was funded, in part, by the National Science Foundation, Schmidt Sciences, the National Bureau of Economic Research, and Columbia University.