先前的工作表明,上下文演示可以越狱语言模型,,但仍不清楚模型如何解释不同类型的合规性演示。我们通过将良性合规性演示(无害请求,有用响应)与有害合规性演示(有害请求,有用响应)混合并测试关于演示组合如何驱动有害合规性的三个假设来研究这一点。在四个模型中, 我们发现良性演示和有害演示不可互换: 良性演示可以减少或增加有害的合规性,具体取决于模型。我们进一步表明,偏好优化是关键的训练阶段,可以防止良性演示增加有害的依从性,,演示排序表现出强烈的新近偏差,,并且模型在拒绝与上下文学习的交互方式上有所不同:,一些模型即使在拒绝时也采用演示的格式,,而另一些则在拒绝时覆盖所有上下文信号。总而言之, 这项工作不仅展示了基于演示的越狱工作,还描述了其工作原理: 从合规性演示中提取的模型取决于演示内容, 排序, 和培训方法。

Prior work has shown that in-context demonstrations can jailbreak language models, but it remains unclear how models interpret different types of compliance demonstrations. We study this by mixing benign compliance demonstrations (non-harmful request, helpful response) with harmful compliance demonstrations (harmful request, helpful response) and testing three hypotheses about how demonstration composition drives harmful compliance. Across four models, we find that benign and harmful demonstrations are not interchangeable: benign demonstrations can either reduce or increase harmful compliance depending on the model. We further show that preference optimization is the critical training stage that prevents benign demonstrations from increasing harmful compliance, that demonstration ordering exhibits strong recency bias, and that models differ in how refusal interacts with in-context learning: some adopt demonstrated formatting even when refusing, while others override all in-context signals upon refusal. Taken together, this work moves beyond showing that demonstration-based jailbreaking works to characterizing how it works: what models extract from compliance demonstrations depends on demonstration content, ordering, and training methodology.

科目: 人工智能 (cs.AI); 机器学习 (cs.LG)

Subjects: Artificial Intelligence (cs.AI); Machine Learning (cs.LG)