视觉语言模型 (VLMs) 很难在交互式物理推理, 中进行泛化,特别是在看不见的任务和环境下。两个关键的故障模式是突出的:幻觉思想链(CoT)推理与物理现实相矛盾,以及模型的推理和行动之间的不一致。我们提出了 VAORA(Visual Action Outcome Outcome Reasoning Alignment), 一种新颖的奖励设计,可以直接解决这两个问题。 VAORA 引入了两种互补奖励: 视觉对齐奖励,,将 VLM 推理锚定到独立于代理动作本身, 的视觉上下文;以及视觉动作对齐奖励,,将推理基于模型的 动作引发的视觉结果。 , 这些奖励共同抑制了幻觉 CoT,并缩小了推理与行为之间的差距。为了提高训练稳定性,,我们通过使用预先训练的领域内专家代理来估计成功概率,进一步采用 smooth, 密集奖励。 PHYRE 和虚拟工具上的实验支持我们在新任务和未见过的环境设置中的表现,,证实可以通过 VAORA 诱导扎根且可推广的物理智能。
Vision-language models (VLMs) struggle to generalize in interactive physical reasoning, particularly under unseen tasks and environments. Two key failure modes are prominent: hallucinated chain-of-thought (CoT) reasoning that contradicts physical reality, and misalignment between the model的 reasoning and actions. We present VAORA (Visual Action Outcome Reasoning Alignment), a novel reward design that directly addresses both issues. VAORA introduces two complementary rewards: Visual Alignment Reward, which anchors VLM reasoning to the visual context independent of the agent action itself, and Visual-Action Alignment Reward, which grounds reasoning in the visual outcome induced by the model的 action. Together, these rewards suppress hallucinated CoT and reduce the gap between reasoning and behavior. To improve training stability, we further employ smooth, dense rewards by estimating success probabilities using a pre-trained in-domain expert agent. Experiments on PHYRE and Virtual Tool support our performances across novel-task and unseen-environment settings, confirming that grounded and generalizable physical intelligence can be induced through VAORA.
科目: 人工智能 (cs.AI); 计算机视觉和模式识别 (cs.CV)
Subjects: Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)