能够自主规划和扩展环境交互的代理人工智能系统提出了一个基本的控制问题:人类如何对可能超出其自身能力的系统进行有意义的监督?现有的可扩展监督方法依赖于复杂的假设,仍然很大程度上是启发式的,或缺乏具有统计保证的顺序设置的实用方法。我们引入了校准集体监督(CCO),,它将不同的辅助评分函数聚合成一个惩罚,衡量与保守基线的偏差。受可实现的效用保护, CCO 的启发,集体保守主义: 行动面临与监督者关注度成正比的惩罚, 因此,当监督者发现高效用行动无异议时,仍然会选择它们,并且只有当关注度累积时才会被推翻。 CCO 使用保形决策理论, 在线校准这种保守性,确保不良结果保持在用户指定的目标阈值以下,且具有有限时间界限且无分布假设。在 SWE-bench, 的修改版本上,较弱的监督者成功地限制了 MACHIAVELLI, CCO 上的对抗性错位的较强代理;,从而大大减少了道德违规,同时保留了奖励。在这两种设置中,, 经验违规率与理论预测的指定目标, 非常接近。

Agentic AI systems capable of autonomous planning and extended environmental interaction pose a fundamental control problem: how can humans maintain meaningful oversight of systems that may exceed their own capabilities? Existing approaches to scalable oversight rely on complex assumptions, remain largely heuristic, or lack practical methods for sequential settings with statistical guarantees. We introduce Calibrated Collective Oversight (CCO), which aggregates diverse auxiliary scoring functions into a penalty measuring deviation from a conservative baseline. Inspired by Attainable Utility Preservation, CCO enables collective conservatism: actions face a penalty proportional to overseer concern, so high-utility actions are still selected when overseers find them unobjectionable and overridden only when concern accumulates. CCO calibrates this conservatism online using Conformal Decision Theory, ensuring that undesirable outcomes remain below a user-specified target threshold with finite-time bounds and no distributional assumptions. On a modified version of SWE-bench, weaker overseers successfully constrain an adversarially misaligned stronger agent; on MACHIAVELLI, CCO substantially reduces ethical violations while preserving reward. In both settings, empirical violation rates closely match the specified targets, as predicted by the theory.

科目:人工智能(cs.AI)

Subjects: Artificial Intelligence (cs.AI)