随着推理能力和部署范围同步增长,大型语言模型(LLMs)获得参与服务于其自身目标的行为的能力,我们将一类风险称为紧急战略推理风险(ESRRs)。其中包括,但不限于,欺骗(故意误导用户或评估者),评估游戏(在安全测试期间有策略地操纵性能),和奖励黑客(利用错误指定的目标)。系统地理解和衡量这些风险仍然是一个开放的挑战。为了解决这一差距,,我们引入了 ESRRSim,,这是一个分类驱动的代理框架,用于自动行为风险评估。我们构建了 7 个类别, 的可扩展风险分类法,该分类法被分解为 20 个子类别。 ESRRSim 生成评估场景,旨在引出忠实的推理,,并搭配双评分标准,在与判断无关且可扩展的架构中评估模型响应和推理痕迹,。对 11 个推理法学硕士的评估显示,风险状况(检测率存在显着差异,范围为 14.45%-72.72%),,随着一代代的显着改进,表明模型可能会越来越多地识别和适应评估环境。
As reasoning capacity and deployment scope grow in tandem, large language models (LLMs) gain the capacity to engage in behaviors that serve their own objectives, a class of risks we term Emergent Strategic Reasoning Risks (ESRRs). These include, but are not limited to, deception (intentionally misleading users or evaluators), evaluation gaming (strategically manipulating performance during safety testing), and reward hacking (exploiting misspecified objectives). Systematically understanding and benchmarking these risks remains an open challenge. To address this gap, we introduce ESRRSim, a taxonomy-driven agentic framework for automated behavioral risk evaluation. We construct an extensible risk taxonomy of 7 categories, which is decomposed into 20 subcategories. ESRRSim generates evaluation scenarios designed to elicit faithful reasoning, paired with dual rubrics assessing both model responses and reasoning traces, in a judge-agnostic and scalable architecture. Evaluation across 11 reasoning LLMs reveals substantial variation in risk profiles (detection rates ranging 14.45%-72.72%), with dramatic generational improvements suggesting models may increasingly recognize and adapt to evaluation contexts.
科目:人工智能(cs.AI)
Subjects: Artificial Intelligence (cs.AI)