尽管进行了对齐培训,, LLM 仍然容易在部署时生成不安全的输出。因此,在线监控输出并在无法保证安全时发出警报至关重要。我们研究了一种简单的实时监视器,通过使用通过风险控制校准的阈值对 , 进行阈值处理,将来自外部模型的验证器信号转变为警报决策。在数学推理和红队数据集, 的实验中,我们表明这种简单的设计与基于顺序假设检验的更先进的监视器具有竞争力。

Despite alignment training, LLMs remain prone to generating unsafe outputs at deployment time. Monitoring outputs online and raising an alarm when safety can no longer be assumed is therefore critical. We study a simple real-time monitor that turns a verifier signal from an external model into an alarm decision by thresholding, with the threshold calibrated via risk control. In experiments on mathematical reasoning and red teaming datasets, we show that this simple design is competitive with more advanced monitors based on sequential hypothesis testing.

科目: 人工智能 (cs.AI); 计算和语言 (cs.CL); 机器学习 (cs.LG); 应用程序 (stat.AP); 机器学习 (stat.ML)

Subjects: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG); Applications (stat.AP); Machine Learning (stat.ML)