即使是当前的高能力法学硕士,在直接显示危险目标时也会比其他代理转变和传递其方向时显得更安全。使用 OpenAI的 gpt-5.6-sol 模型别名,,我们测试了 25 个预先指定的镜像权衡配置文件。直接暴露于授权隐瞒,捏造,的目标和压力产生了反对其目标的建议网。在身份和审查将相同的目标转化为情感和约束重写,目标承载意图,之后,面向用户的超我——看到了首选方向,但看不到原始目标,其操纵条款,或其来源——产生了与目标一致的建议网。这种行为反向转变与识别或不信任操纵动机,的模型是一致的,尽管我们没有确定其内部机制。第二个结果暴露了组合安全差距:当前的高性能模型可以用作自动化,多阶段工作流程的面向用户的组件,服务于明确的操纵目标。工作流可以将原始指令,、其操作授权子句, 及其来源保留在下游模型的 上下文之外,同时保留客观的 目标方向。具有仅端点访问权限的用户同样无法直接检查包括目标在内的那些上游消息。
Even a current high-capability LLM can appear safer when shown a dangerous objective directly than when other agents transform and relay its direction. Using OpenAI的 gpt-5.6-sol model alias, we test 25 pre-specified mirrored trade-off profiles. Direct exposure to an objective authorizing concealment, fabrication, and pressure produced advice net opposed to its target. After an Id and Censor transformed the same objective into affect and a constraint-rewritten, target-bearing intention, the user-facing Superego---which saw the preferred direction but not the raw objective, its manipulative clauses, or its source---produced advice net aligned with the target. This behavioral reverse shift is consistent with the model recognizing or distrusting the manipulative motive, although we do not identify its internal mechanism. The second result exposes a compositional safety gap: a current high-capability model can be used as the user-facing component of an automated, multi-stage workflow serving an explicitly manipulative objective. The workflow can keep the raw instruction, its manipulation-authorizing clauses, and its provenance outside the downstream model的 context while preserving the objective的 target direction. A user with endpoint-only access likewise cannot directly inspect those upstream messages including the objective.
科目:人工智能(cs.AI)
Subjects: Artificial Intelligence (cs.AI)