人工智能系统正在进入医疗保健,、金融,和国防,等关键领域,但仍然容易受到对抗性攻击。虽然 AI 红队是主要防御,,但当前的方法迫使操作员进入手动, 特定于库的工作流程。操作员花费数周时间手工制作工作流程 - 组装攻击, 转换, 和记分器。当结果达不到, 时,必须重建工作流程。因此,, 操作员花费更多时间构建工作流程,而不是探测目标的安全性和安全漏洞。我们引入了一个基于开源 Dreadnode SDK 构建的 AI 红队代理。该代理创建基于 45+ 对抗性攻击,、450+ 转换, 和 130+ 记分器的工作流程。操作员可以探测多代理系统,、多语言, 和多模式目标,,重点关注探测内容而不是如何实施。我们做出三个贡献: 1.代理界面。操作员通过 Dreadnode TUI (终端用户界面) 以自然语言描述目标。该代理处理攻击选择, 转换组合, 执行, 和报告, 让操作员专注于红队。几周压缩为几个小时。 2.统一的框架。用于探测传统 ML 模型 ( 对抗性示例 ) 和生成 AI 系统 ( 越狱), 的单一框架,无需单独的库。 3. Llama Scout 案例研究。我们红队 Meta Llama Scout,使用零人类开发的代码实现了 85% 的攻击成功率,严重性高达 1.0,
AI systems are entering critical domains like healthcare, finance, and defense, yet remain vulnerable to adversarial attacks. While AI red teaming is a primary defense, current approaches force operators into manual, library-specific workflows. Operators spend weeks hand-crafting workflows - assembling attacks, transforms, and scorers. When results fall short, workflows must be rebuilt. As a result, operators spend more time constructing workflows than probing targets for security and safety vulnerabilities. We introduce an AI red teaming agent built on the open-source Dreadnode SDK. The agent creates workflows grounded in 45+ adversarial attacks, 450+ transforms, and 130+ scorers. Operators can probe multi-agent systems, multilingual, and multimodal targets, focusing on what to probe rather than how to implement it. We make three contributions: 1. Agentic interface. Operators describe goals in natural language via the Dreadnode TUI (Terminal User Interface). The agent handles attack selection, transform composition, execution, and reporting, letting operators focus on red teaming. Weeks compress to hours. 2. Unified framework. A single framework for probing traditional ML models (adversarial examples) and generative AI systems (jailbreaks), removing the need for separate libraries. 3. Llama Scout case study. We red team Meta Llama Scout and achieve an 85% attack success rate with severity up to 1.0, using zero human-developed code
科目: 人工智能 (cs.AI); 密码学和安全 (cs.CR)
Subjects: Artificial Intelligence (cs.AI); Cryptography and Security (cs.CR)