当大型语言模型 (LLMs) 无法泛化或在推理, 时出现偶然错误时,通常会被视为 LLM 并未真正推理, 而是执行某种模式匹配的证据。这意味着人们的行为不会表现出相同类型的失败,因为人类推理使用原则性和抽象的世界模型。我们评估了人类参与者和 25 名法学硕士对各种日常情况进行常识推理的能力,并观察了人和模型中类似的错误模式。然后,我们识别驱动 LLM 响应的一组注意力头,并发现这些头实现了某种形式的模式匹配。这些注意力头使我们能够预测人们由于表面上不相关的提示细节而导致的看似无法解释的推理错误。综合起来,,我们的结果表明,人们和法学硕士的日常因果推理与模式匹配的形式比与抽象世界模型更一致。
When large language models (LLMs) fail to generalize or make haphazard errors in reasoning, it is often taken as evidence that LLMs are not truly reasoning, but rather performing a kind of pattern matching. The implication is that people的 behavior does not exhibit the same types of failures because human reasoning uses principled and abstract world models. We evaluate human participants and 25 LLMs on their ability to engage in common-sense reasoning about a variety of everyday situations and observe similar patterns of errors in both people and models. We then identify the set of attention heads driving LLM responses and find that these heads implement a form of pattern-matching. These attention heads allow us to predict seemingly inexplicable reasoning errors in people caused by ostensibly irrelevant prompt details. Taken together, our results suggest that everyday causal reasoning in people and LLMs is more consistent with a form of pattern-matching than with abstract world models.
科目:人工智能(cs.AI)
Subjects: Artificial Intelligence (cs.AI)