多模态大语言模型 (MLLMs) 越来越多地用于解释可视化,,但当前的评估仍然主要以图表为中心,并提供有限的科学可视化理解证据(SciVis)。我们在科学可视化素养评估测试, 上对 6 个 MLLM 进行了基准测试,这是一种标准化的 SciVis 素养评估,包括基于 18 种科学可视化和插图的 49 个项目,,涵盖 8 种技术和 11 种任务类型。我们在封闭世界协议下评估了三个闭源模型和三个开源模型,并使用 485 名人类参与者的数据比较了它们的性能。结果表明,当前的 MLLM 并未表现出统一的科学素养。 Gemini 是整体最强的模型,,在评估的子集, 中超过了人类平均值,而开源模型仍低于人类基线。不同技术和任务的性能非常不平衡: 模型在科学插图, 搜索, 和空间理解, 方面表现最好,但在基于纹理和基于集成的可视化以及定量估计方面表现不佳。误差分析揭示了细粒度定量估计,流向解释,和接地编码解释中反复出现的失败。这些发现将 SciVis 素养定位为评估多模式人工智能系统的必要基准维度。我们的代码和模型输出可通过此 https URL 公开获得。
Multimodal large language models (MLLMs) are increasingly used to interpret visualizations, yet current evaluations remain largely chart-centric and provide limited evidence of understanding of scientific visualization (SciVis). We benchmark six MLLMs on the scientific visualization literacy assessment test, a standardized SciVis literacy assessment comprising 49 items based on 18 scientific visualizations and illustrations, spanning 8 techniques and 11 task types. We evaluate three closed-source and three open-source models under a closed-world protocol and compare their performance using data from 485 human participants. Results show that current MLLMs do not exhibit uniform SciVis literacy. Gemini is the strongest model overall, exceeding the human mean across the evaluated subsets, whereas the open-source models remain below the human baseline. Performance is highly uneven across techniques and tasks: models perform best on scientific illustration, search, and spatial understanding, but struggle on texture-based and integration-based visualizations and on quantitative estimation. Error analysis reveals recurring failures in fine-grained quantitative estimation, flow-direction interpretation, and grounded encoding interpretation. These findings position SciVis literacy as a necessary benchmark dimension for evaluating multimodal AI systems. Our code and model outputs are publicly available at this https URL.
科目: 人工智能 (cs.AI); 计算与语言 (cs.CL); 人机交互 (cs.HC)
Subjects: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Human-Computer Interaction (cs.HC)