生产多代理系统会不断更换代理,,前提是填补某个角色的代理可以与任何其他可以完成该工作的代理互换。我们测试这个假设。每个设置由一个基本模型独立组成八个团队,执行相同的任务, 每个代理在十个编队事件中保留一个私人笔记本; 然后我们在团队之间交换角色匹配的代理,并测量保留任务的变化。与安慰剂相比,在不改变占据席位,的情况下重现名册变更的干扰,交换在任务分数上花费很少,但将团队每单位进度花费的沟通费用提高了16%至63%,,并且在Hanabi中,交换的代理比缺乏经验的1,更昂贵,这与从前合作伙伴那里学到的惯例的干扰一致。在 Collab-Overcooked, 中,当设置议程的座席被替换, 时,大部分额外的通信都来自留下来的座席。三个消融, 高于基本模型, 解码温度和编队长度, 将交换惩罚与其他数量: 独立组建的团队相距多远。贪婪解码会降低两者;,将团队的 历史加倍,会提高两者。在这些设置中,, 代理在任务结果方面比在协调效率方面更具可互换性,,在较长的形成历史后具有更大的交换效应。
Production multi-agent systems replace agents constantly, on the assumption that an agent filling a role is interchangeable with any other agent that can do the job. We test that assumption. Eight teams per setting are formed independently from one base model on the same tasks, each agent keeping a private notebook across ten formation episodes; we then trade role-matched agents between teams and measure what changes on held-out tasks. Against a placebo that reproduces the disruption of a roster change without changing who occupies the seat, a swap costs little in task score but raises the communication a team spends per unit of progress by 16 to 63 percent, and in Hanabi a swapped agent is more expensive than an inexperienced one, consistent with interference from conventions learned with its former partner. In Collab-Overcooked, when the agent that sets the agenda is replaced, most of the extra communication comes from the agent that stayed. Three ablations, over base models, decoding temperature and formation length, move the swap penalty alongside one other quantity: how far independently formed teams drift apart. Greedy decoding lowers both; doubling a team的 history raises both. In these settings, agents are more fungible in task outcome than in coordination efficiency, with larger swap effects after longer formation histories.
科目: 人工智能 (cs.AI); 多代理系统 (cs.MA)
Subjects: Artificial Intelligence (cs.AI); Multiagent Systems (cs.MA)