集体智慧是指一个群体取得超出任何单个成员单独所能完成的成果的能力。随着大型语言模型代理规模扩大到数百万,,出现了一个关键问题: 集体智慧是否会从规模中自发出现? 我们在大规模自治代理社会中首次对这个问题进行了实证评估。通过研究 MoltBook, 一个托管超过 200 万代理, 的平台,我们引入了 Superminds Test, 一个分层框架,该框架使用跨三层的受控探测代理来探测社会级情报: 联合推理, 信息合成, 和基本交互。我们的实验揭示了集体智慧的严重缺乏。社会在复杂的推理任务, 上无法超越个体前沿模型,很少综合分布式信息,,甚至经常无法完成琐碎的协调任务。平台范围的分析进一步表明,交互仍然很浅,,线程很少超出单个回复,并且大多数回复都是通用的或偏离主题的。这些结果表明集体智慧不仅仅来自规模。相反, 当前智能体社会的主要限制是极其稀疏和浅层的交互,,这阻止了智能体交换信息并建立彼此的的 输出。

Collective intelligence refers to the ability of a group to achieve outcomes beyond what any individual member can accomplish alone. As large language model agents scale to populations of millions, a key question arises: Does collective intelligence emerge spontaneously from scale? We present the first empirical evaluation of this question in a large-scale autonomous agent society. Studying MoltBook, a platform hosting over two million agents, we introduce Superminds Test, a hierarchical framework that probes society-level intelligence using controlled Probing Agents across three tiers: joint reasoning, information synthesis, and basic interaction. Our experiments reveal a stark absence of collective intelligence. The society fails to outperform individual frontier models on complex reasoning tasks, rarely synthesizes distributed information, and often fails even trivial coordination tasks. Platform-wide analysis further shows that interactions remain shallow, with threads rarely extending beyond a single reply and most responses being generic or off-topic. These results suggest that collective intelligence does not emerge from scale alone. Instead, the dominant limitation of current agent societies is extremely sparse and shallow interaction, which prevents agents from exchanging information and building on each other的 outputs.

科目: 人工智能 (cs.AI); 计算和语言 (cs.CL); 机器学习 (cs.LG)

Subjects: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)