LLM 代理的最新进展实现了复杂的认知能力,,例如多步推理, 规划, 和工具使用,,这些能力越来越多地将这些代理定位为人类协作者。然而,有效的协作, 要求协作者在协作过程中不断维护和调整自己推理,合作伙伴 意图, 的心理模型和共同目标。如今,的智能体很少开发此类功能,因为它们主要是针对任务完成, 进行优化的,并且社区缺乏具有动作级心智模型注释的真实人类协作数据,这些数据可以指导智能体实现流程级协作能力。为了弥补这一差距,,我们提出了 ALMANAC, ,这是一个用于代理协作的动作级心智模型 ANnotations 的数据集,它是根据社会科学中的经典二元路由任务 Map Task, 构建的。年鉴包含 2,987 个协作操作,,每个操作都与理论指导的心理模型注释配对,记录参与者 自我推理, 感知的合作伙伴意图, 和感知的团队目标。我们对 6 个法学硕士进行了预测人类 下一回合行为和心理模型的基准测试。我们的结果证明了 ALMANAC的 在评估模型 模拟人类协作行为并推断其潜在心理模型的能力方面的实用性。
Recent advances in LLM agents have enabled complex cognitive capabilities, such as multi-step reasoning, planning, and tool use, that increasingly position these agents as human collaborators. Effective collaboration, however, requires collaborators to continuously maintain and align mental models of their own reasoning,partners intentions, and shared goals during the collaborative process. Today的 agents rarely develop such capabilities since they are primarily optimized for task completion, and the community lacks authentic human collaboration data with action-level mental model annotations that could guide agents toward process-level collaborative competence. To bridge this gap, we present ALMANAC, a dataset of Action-Level Mental model ANnotations for Agent Collaboration built from the Map Task, a classic dyadic routing task from social science. ALMANAC contains 2,987 collaboration actions, each paired with theory-informed mental model annotations that record the participants self-reasoning, perceived partner intent, and perceived team goal. We benchmark six LLMs on predicting humans next-turn behavior and mental models. Our results demonstrate ALMANAC的 utility in evaluating models ability to simulate human collaborative behaviors and infer their underlying mental models.
科目: 人工智能 (cs.AI); 计算和语言 (cs.CL)
Subjects: Artificial Intelligence (cs.AI); Computation and Language (cs.CL)