自私的代理人,不受约束,倾向于在重复的社会困境中叛逃,,导致贸易合作收益崩溃。本文研究了在不受限制的通信, 之上分层的正式机制, 足以让此类代理的社会维持市场稳定性, 以及这些机制对对抗性攻击的弹性如何。我们将研究问题实例化为多代理市场模拟,其中具有互补生产专业的 18 个 LLM 代理 (DeepSeek-V3) 必须在受限的社交网络内进行交易以获得效用。我们进行了两个实验阶段: (1) 在超过 200 轮的渐进巨魔注入下的八种条件下进行机制比较, 将调解确定为表现最佳的机制; 和 (2) 使用迭代提示优化的 LLM 驱动的巨魔进行调解的对抗性红队, 发现最佳攻击 (v6) 通过以下方式降低诚实代理效用13.3% 但不能让市场崩溃。即使在持续的敌对压力下,调解也能实现恢复。我们将对抗鲁棒性定义为一种机制的 在优化攻击下维持积极诚实代理效用的能力,,并发现中介是鲁棒的: 它可以弯曲但不能破坏。
Self-interested agents, left unconstrained, tend toward defection in repeated social dilemmas, causing cooperative gains from trade to collapse. This paper investigates what formal mechanisms, layered on top of unrestricted communication, are sufficient for a society of such agents to maintain market stability, and how resilient those mechanisms are to adversarial attack. We instantiate the research question as a multi-agent marketplace simulation where 18 LLM agents (DeepSeek-V3) with complementary production specialties must trade within a constrained social network to obtain utility. We conduct two experimental phases: (1) a mechanism comparison across eight conditions under progressive troll injection over 200 rounds, identifying Mediation as the top-performing mechanism; and (2) adversarial red-teaming of Mediation using iteratively prompt-optimised LLM-driven trolls, finding that the best attack (v6) reduces honest-agent utility by 13.3% but cannot collapse the market. Mediation enables recovery even under sustained adversarial pressure. We define adversarial robustness as a mechanism的 ability to sustain positive honest-agent utility under optimised attack, and find that Mediation is robust: it can be bent but not broken.
科目:人工智能(cs.AI)
Subjects: Artificial Intelligence (cs.AI)