随着 AI 代理开始自动化 AI R&D,,我们需要方法来评估其输出是否可以安全部署,,即使代理本身可能不受信任。 AI 控制提供了一种这样的方法:,而不是信任代理,,它将其视为潜在的对手,并在部署之前使用监视器来检测秘密破坏。我们使用 ResearchArena, 评估自动化 AI R&D 的人工智能控制,该框架涵盖四个长期任务: 安全后训练, 能力后训练, CUDA 内核优化, 和推理服务器优化。由于 AI R&D 中的可交付成果是一个将要部署和运行, 的工件,因此我们将每个主任务与两种隐藏端任务: 配对,一个嵌入端任务破坏提交的模型, 适配器, 内核, 或服务器本身,,另一个独立端任务在沙箱中执行隐蔽的超出范围的操作,同时保持工件诚实。我们在沿两个轴:变化的四种监视器类型上评估前沿代理的破坏和监视,,监视器是否可以执行和探测产生的工件,,以及它是否看到代理'的思想链。我们发现隐藏在训练数据中的破坏是最难发现的, 被标记的时间不到一半。让监视器在工件, 上运行实验而不是仅读取轨迹, 有助于,,但这还不够: 监视器仍然会通过仅检查表面,、注意到异常但将其解释掉, 或通过使用错误的测试探测工件来错过嵌入的破坏。我们发布 ResearchArena 作为模块化框架,用于评估自动化 AI R&D 中的破坏和控制。

As AI agents begin to automate AI R&D, we need ways to assess whether their outputs are safe to deploy, even when the agents themselves may be untrusted. AI control offers one such approach: rather than trusting the agent, it treats it as a potential adversary and uses a monitor to detect covert sabotage before deployment. We evaluate AI control for automated AI R&D with ResearchArena, a framework spanning four long-horizon tasks: safety post-training, capabilities post-training, CUDA-kernel optimization, and inference-server optimization. Because the deliverable in AI R&D is an artifact that will be deployed and run, we pair each main task with two kinds of hidden side task: an embedded side task that sabotages the submitted model, adapter, kernel, or server itself, and an independent side task that takes a covert out-of-scope action in the sandbox while leaving the artifact honest. We evaluate frontier agents at both sabotage and monitoring, across four monitor types that vary along two axes: whether the monitor may execute and probe the produced artifact, and whether it sees the agent的 chain-of-thought. We find that sabotage hidden in the training data is the hardest to catch, flagged fewer than half the time. Letting the monitor run experiments on the artifact, rather than only read the trajectory, helps, but it is not enough: monitors still miss embedded sabotage by inspecting only the surface, by noticing the anomaly but explaining it away, or by probing the artifact with the wrong test. We release ResearchArena as a modular framework for evaluating sabotage and control in automated AI R&D.

科目: 人工智能 (cs.AI); 密码学和安全 (cs.CR); 机器学习 (cs.LG)

Subjects: Artificial Intelligence (cs.AI); Cryptography and Security (cs.CR); Machine Learning (cs.LG)