强世界模型代理经常包含弱世界模型。我们通过在 Atari Pong: DreamerV3, DIAMOND, TWISTER, Simulus, 和 STORM, 中重现 5 个视觉世界模型代理来研究这一代理世界模型差距,其性能与报告的结果, 相当,并独立评估他们的冻结世界模型。 First, 闭环推出诊断定性检查每个冻结模型在独立训练的策略下生成的视觉轨迹。所有五个模型都表现出明显的视觉或动态故障,,包括球消失, 不正确的运动, 和无效的球-桨相互作用。第二, 在基于本机零样本模型的强化学习(MBRL), 下,使用代理的 本机RL 过程, 完全在冻结模型中从头开始训练新策略,而无需进行真实环境训练。在真实环境, 中进行评估时,这些策略的性能明显低于复制代理: DreamerV3 (-5.5 至 -20.9), DIAMOND (19.7 至 -9.6), TWISTER (17.7 至 -13.3), Simulus (20.8 至 -11.6), 和STORM (18.7 至 -21.0),,其中 -21 是最小 Pong 回报。 Atari100K 也存在这种差距。受 Pong, 中与球相关的推出失败的启发,我们提出概念引导空间正则化(CGSReg),,这是任务关键概念区域的辅助重建损失。我们在更具挑战性的像素空间零样本 MBRL 设置, 下对其进行评估,其中策略直接从冻结世界模型生成的图像中学习。 Ball-region CGSReg 改进了 DreamerV3, DIAMOND, TWISTER, 和 Simulus, 中的像素空间零样本 MBRL,并且还改进了前三个; STORM 中的闭环推出,但没有显示明显的改进。

Strong world-model agents frequently contain weak world models. We study this agent-world-model gap by reproducing five visual world-model agents in Atari Pong: DreamerV3, DIAMOND, TWISTER, Simulus, and STORM, with performance comparable to the reported results, and independently evaluating their frozen world models. First, closed-loop rollout diagnosis qualitatively inspects visual trajectories generated by each frozen model under an independently trained policy. All five models exhibit clear visual or dynamical failures, including ball disappearance, incorrect motion, and invalid ball-paddle interactions. Second, under native zero-shot model-based reinforcement learning (MBRL), a new policy is trained entirely within the frozen model from scratch using the agent的 native RL procedure, without real-environment training. When evaluated in the real environment, these policies substantially underperform the reproduced agents: DreamerV3 (-5.5 to -20.9), DIAMOND (19.7 to -9.6), TWISTER (17.7 to -13.3), Simulus (20.8 to -11.6), and STORM (18.7 to -21.0), where -21 is the minimum Pong return. This gap also extends broadly across Atari100K. Motivated by the ball-related rollout failures in Pong, we propose Concept-Guided Spatial Regularization (CGSReg), an auxiliary reconstruction loss on task-critical concept regions. We evaluate it under a more challenging pixel-space zero-shot MBRL setting, where policies learn directly from images generated by the frozen world model. Ball-region CGSReg improves pixel-space zero-shot MBRL in DreamerV3, DIAMOND, TWISTER, and Simulus, and also improves closed-loop rollouts in the first three; STORM shows no clear improvement.

科目: 人工智能 (cs.AI); 机器学习 (cs.LG)

Subjects: Artificial Intelligence (cs.AI); Machine Learning (cs.LG)