人类在遇到新环境时会快速学习抽象知识,并灵活运用这些知识来指导高效、智能的行动。现代人工智能系统能否以类似的方式学习和计划? 我们使用复杂的人类游戏数据集和并发 fMRI 记录, 研究这个问题,其中参与者学习需要规则发现, 假设修正, 和多步骤规划的新颖视频游戏。我们通过玩游戏,匹配人类学习行为,并在同一任务,期间预测大脑活动的能力来共同评估模型,将一套前沿大型推理模型(LRMs)与无模型和基于模型的深度强化学习代理和基于贝叶斯理论的代理进行比较。我们发现,前沿 LRM 在游戏发现过程中与人类行为模式最接近,并且预测大脑活动的数量级比皮质和皮质下区域的强化学习替代方案好一个数量级,,且对排列控制具有稳健的效果。通过有针对性的操作,,我们进一步表明大脑对齐反映了模型的 游戏状态的上下文表示,而不是其下游计划或推理。我们的结果将 LRM 确立为复杂, 自然环境中人类学习和决策的引人注目的计算帐户。具有交互式重播的项目页面: 此 https URL
Humans rapidly learn abstract knowledge when encountering novel environments and flexibly deploy this knowledge to guide efficient and intelligent action. Can modern AI systems learn and plan in a similar way? We study this question using a dataset of complex human gameplay with concurrent fMRI recordings, in which participants learn novel video games that require rule discovery, hypothesis revision, and multi-step planning. We jointly evaluate models by their ability to play the games, match human learning behavior, and predict brain activity during the same task, comparing a suite of frontier Large Reasoning Models (LRMs) against model-free and model-based deep reinforcement learning agents and a Bayesian theory-based agent. We find that frontier LRMs most closely match human behavioral patterns during game discovery and predict brain activity an order of magnitude better than both reinforcement learning alternatives across cortical and subcortical regions, with effects robust to permutation controls. Through targeted manipulations, we further show that brain alignment reflects the model的 in-context representation of the game state rather than its downstream planning or reasoning. Our results establish LRMs as compelling computational accounts of human learning and decision making in complex, naturalistic environments. Project page with interactive replays: this https URL
科目: 人工智能 (cs.AI); 神经元和认知 (q-bio.NC)
Subjects: Artificial Intelligence (cs.AI); Neurons and Cognition (q-bio.NC)