与表现出强大推理能力的大型语言模型 (LLMs) 不同,, 视觉语言模型 (VLMs) 很难进行视觉推理,,即使在允许等效文本, 图, 和组合图+ 文本视图的几何问题上也是如此。我们表明,这些视图通常会引发不同的行为: 模型可能通过文本解决问题,但在相应的图表上失败, 或者在视觉上成功而在文本上失败。这种不一致表明,不同的观点暴露了标准多模式后训练没有充分利用的互补推理路径和故障模式。为了研究和利用这种现象,,我们构建了 ODA-Data, 一个高质量的配对多模态几何数据集,其中以文本为主, 图像为主, 和相同问题, 的组合图像+文本视图以及用于训练和评估模态依赖推理行为的分割。然后,我们开发了模态通知交互推理优化(MIRROR),,这是一种通过自我监督改进多模态推理的强化学习方法。对于每个问题, MIRROR 在所有视图下评估模型, 选择表现最好的视图作为教师, 并以针对教师的反向KL 目标训练其他视图。在评估几何问题的推理基准中, MIRROR 比标准 RL 有所改进,并在不同模式下产生更准确和一致的行为
Unlike large language models (LLMs) that exhibit strong reasoning capabilities, vision-language models (VLMs) struggle with visual reasoning, even on geometry problems that admit equivalent text, diagram, and combined diagram+text views. We show that these views often elicit different behaviors: a model may solve a problem from text but fail on the corresponding diagram, or succeed visually while failing textually. This inconsistency suggests that different views expose complementary reasoning paths and failure modes that standard multimodal post-training does not fully exploit. To study and exploit this phenomenon, we construct ODA-Data, a high-quality paired multimodal geometry dataset with text-dominant, image-dominant, and combined image+text views of the same problems, together with splits for training and evaluating modality-dependent reasoning behaviors. We then develop Modality-Informed Reciprocal Reasoning Optimization (MIRROR), a reinforcement learning approach for improving multimodal reasoning via self supervision. For each problem, MIRROR evaluates the model under all views, selects the best-performing view as a teacher, and trains other views with a reverse-KL objective towards the teacher. Across reasoning benchmarks that evaluate on geometry problems, MIRROR improves over standard RL and yields more accurate and consistent behavior across modalities
科目: 人工智能 (cs.AI); 机器学习 (cs.LG)
Subjects: Artificial Intelligence (cs.AI); Machine Learning (cs.LG)