"thinking-with-images"范式为多模式法学硕士配备了主动视觉操作,例如裁剪和缩放。然而,使用这些操作的 , 模型通常仅以比直接推理更高的代币成本获得边际收益或负收益。他们还可能会反复裁剪不相关的区域,并且无法正确回答直接推理答案的问题。我们询问返回的视觉证据是否会因果影响答案。为了回答这个问题,,我们将视觉工具的使用制定为因果图,将观察介导的路径与行动引起的捷径分开。然后,我们通过三个级别的干预进行审计:政策(将工具使用与直接推理进行比较),轨迹(在推出期间破坏所有观察),和步骤(反事实地替换固定前缀下的单个观察)。我们的步级估计, 视觉证据增益, 隔离了每个返回观察值的贡献。通过六个代表性模型和五个细粒度感知基准,,我们发现了具有两种故障模式的政策校准错误。在 Calling Without Look, 中,返回的观察结果对答案没有因果影响。在“没有计划的观察”中,, 的观察结果提供了丰富的信息,但呼叫时间表不连贯。轨迹级诊断分解了策略级准确性增益,并显示增益集中在校准的少数群体中。我们将这种差异称为视觉工具使用: 的错觉,尽管总体准确度提高了, 视觉工具的使用在广泛的推广中并不具有因果关系。该代码可从此 https URL 获取。
The "thinking-with-images" paradigm equips multimodal LLMs with active visual operations such as crop-and-zoom. However, models using these operations often achieve only marginal or negative gains over direct inference at substantially higher token cost. They may also repeatedly crop irrelevant regions and fail on questions that direct inference answers correctly. We ask whether the returned visual evidence causally affects the answer. To answer this question, we formulate visual tool-use as a causal graph that separates observation-mediated paths from action-induced shortcuts. We then audit it through interventions at the three levels: policy (comparing tool-use with direct inference), trajectory (corrupting all observations during rollout), and step (counterfactually replacing one individual observation under a fixed prefix). Our step-level estimand, Visual Evidence Gain, isolates the contribution of each returned observation. Across six representative models and five fine-grained perception benchmarks, we uncover policy miscalibration with two failure modes. In Calling Without Looking, returned observations have no causal effect on the answer. In Looking Without Planning, observations are informative but the call schedule is incoherent. A trajectory-level diagnostic decomposes the policy-level accuracy gain and shows that the gain is concentrated in a Calibrated minority. We term this discrepancy the illusion of visual tool-use: despite aggregate accuracy gains, visual tool-use is not causally effective across a broad range of rollouts. The code is available at this https URL.
科目:人工智能(cs.AI)
Subjects: Artificial Intelligence (cs.AI)