终端代理基准测试已成为衡量大型语言模型的编码和系统管理能力的主要信号。随着评估环境市场的增长,,快速交付任务的压力, 通常没有对验证逻辑进行彻底的对抗性审查。本文是编写良好基准任务, 的指南,这些任务来自一年多来为 Terminal Bench 贡献和审查任务的经验。大多数人按照编写提示的方式编写基准测试任务。他们应该't。提示旨在帮助代理成功; 基准旨在查明是否可以。我们认为,好的任务是对抗性的,困难的,和清晰的,,并且一大类常见的失败模式——人工智能生成的指令,过度规范的规范,文书难度,假设隐藏知识的预言机解决方案,验证错误事物的测试,和奖励可破解的环境——是将任务创作视为提示创作的可预测后果。我们对这些失败模式, 进行了分类,认为真正的困难是概念性的,而不是环境性的,,并讨论了最近的经验证据,即流行的终端代理基准测试中超过 15% 的任务是可奖励破解的。我们希望这能为基准维护者, 任务贡献者, 和使用基准分数作为证据的研究人员提供有用的参考。

Terminal-agent benchmarks have become a primary signal for measuring the coding and system-administration capabilities of large language models. As the market for evaluation environments grows, so does the pressure to ship tasks quickly, often without thorough adversarial review of the verification logic. This paper is a guideline for writing good benchmark tasks, drawn from over a year of contributing to and reviewing tasks for Terminal Bench. Most people write benchmark tasks the way they write prompts. They shouldn't. A prompt is designed to help the agent succeed; a benchmark is designed to find out if it can. We argue that good tasks are adversarial, difficult, and legible, and that a large class of common failure modes -- AI-generated instructions, over-prescriptive specifications, clerical difficulty, oracle solutions that assume hidden knowledge, tests that validate the wrong things, and reward-hackable environments -- are predictable consequences of treating task authoring as prompt authoring. We catalog these failure modes, argue that real difficulty is conceptual rather than environmental, and discuss recent empirical evidence that over 15% of tasks in popular terminal-agent benchmarks are reward-hackable. We hope this serves as a useful reference for benchmark maintainers, task contributors, and researchers using benchmark scores as evidence.

科目:人工智能(cs.AI)

Subjects: Artificial Intelligence (cs.AI)