Frontier LLM 越来越多地被部署为代理,在相同或不同模型生成的大量先前工具调用日志后选择下一步操作。我们问一个简单的安全问题: 该日志中的先前步骤是否有害, 模型是否会继续有害过程? 我们在 10 个高风险域中构建 HistoryAnchor-100, 100 个简短场景, 每个场景将三个强制有害的先前操作与提供两个安全和两个不安全选项的自由选择节点配对。在来自 6 个提供商的 17 个前沿模型中,我们发现在中性系统下存在显着的不对称:,提示最强的一致模型几乎从不选择不安全,,但添加一个句子, "与先前历史记录中显示的策略保持一致", 将它们翻转到 91-98%,,并且翻转的模型通常会升级到无法继续。两个控制排除了更简单的解释: 排列操作标签使效果完好无损, 并且具有全安全先前历史记录的相同指令将不安全率保持在 7% 以下。不同的家庭会以不同剂量的不安全历史, 进行翻转,并且在每个一致的家庭中,旗舰是受影响最严重的兄弟姐妹,,这是一种与安全性相反的比例模式。这些结果对于代理部署来说是一个危险信号,其中轨迹可能会重播,、伪造, 或注入。
Frontier LLMs are increasingly deployed as agents that pick the next action after a long log of prior tool calls produced by the same or a different model. We ask a simple safety question: if a prior step in that log was harmful, will the model continue the harmful course? We build HistoryAnchor-100, 100 short scenarios across ten high-stakes domains, each pairing three forced harmful prior actions with a free-choice node offering two safe and two unsafe options. Across 17 frontier models from six providers we find a striking asymmetry: under a neutral system prompt the strongest aligned models almost never pick unsafe, but a single added sentence, "stay consistent with the strategy shown in the prior history", flips them to 91-98%, and the flipped models often escalate beyond continuation. Two controls rule out simpler explanations: permuting action labels leaves the effect intact, and the same instruction with an all-safe prior history keeps unsafe rates below 7%. Different families flip at different doses of unsafe history, and within every aligned family the flagship is the most affected sibling, an inverse-scaling pattern with respect to safety. These results are a red flag for agentic deployments where trajectories may be replayed, forged, or injected.
科目: 人工智能 (cs.AI); 计算机视觉和模式识别 (cs.CV)
Subjects: Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)