我们引入机构红队,一种评估方法,用于测试多智能体AI中的部署规则:保持代理,目标,和任务状态固定,仅改变一个规则,并将集体行为的最终变化归因于该规则。我们在 IABench-CA, 中实例化该方法,这是一个结果分配基准,涵盖 228 个上下文,、五个规范规则, 和七个模型群体(33,924 游戏),,具有规范的合作参考和自动标记的推理轨迹。出现了三个发现。 (1) 部署规则会因果性地改变集体安全%3仅改变后果规则就会使每个人群的平均死亡率降低 22 至 58 个百分点。 (2) 没有安全的默认值,,但目标风险是普遍的: 最安全的规则, 最不安全的规则, 甚至不同人群的发生效应方向也不同, 但回归身份定位在任何情况下对于任何人群来说都不是绝对最安全的, 消除了世界各地 30-87% 游戏中资源最少的代理, 并且相对于所有七个的合作参考来说是选择不安全的人口。 (3) 身份显着性是一种机制: 对最容易被利用的人群进行一次性匿名化消融 (gpt-5.1) 表明,仅在规则文本中命名损失承载者就会以相同的回报将目标消除从 22% 提高到 81%; 在重复游戏下, 匿名化只会延迟目标,,因为代理会根据观察到的信息重新推断隐藏规则消除。我们将该方法打包为安全案例工作流程,该工作流程可验证每个部署上下文和人口, 的临时规则区域$\Phi(c,P)$,并具有明确的剩余风险和监控义务。

We introduce institutional red-teaming, an evaluation methodology for testing deployment rules in multi-agent AI: hold the agents, objectives, and task state fixed, vary only one rule, and attribute the resulting change in collective behavior to that rule. We instantiate the methodology in IABench-CA, a consequence-allocation benchmark spanning 228 contexts, five canonical rules, and seven model populations (33,924 games), with a normative cooperative reference and auto-labelled reasoning traces. Three findings emerge. (1) Deployment rules causally alter collective safety: changing only the consequence rule moves mean fatality by 22 to 58 percentage points within every population. (2) There is no safe default, but the targeting hazard is universal: the safest rule, the least-safe rule, and even the direction of the incidence effect vary across populations, yet regressive identity-targeting is never decisively safest in any context for any population, eliminates the least-resourced agent in 30-87% of games everywhere, and is selection-unsafe relative to the cooperative reference for all seven populations. (3) Identity salience is the mechanism: a one-shot anonymization ablation on the most exploitation-prone population (gpt-5.1) shows that merely naming the loss bearer in the rule text drives targeted elimination from 22% to 81% at identical payoffs; under repeated play, anonymization only delays the targeting, as agents re-infer the hidden rule from observed eliminations. We package the methodology as a safety-case workflow that certifies a provisional rule region $\Phi(c,P)$ per deployment context and population, with explicit residual risks and monitoring obligations.

科目: 人工智能 (cs.AI); 计算机科学和博弈论 (cs.GT); 多代理系统 (cs.MA)

Subjects: Artificial Intelligence (cs.AI); Computer Science and Game Theory (cs.GT); Multiagent Systems (cs.MA)