随着 LLM 越来越多地部署在代理系统, 中,它们的能力不仅取决于模型权重,还取决于工具: 提示, 工具, 控制流, 内存, 以及围绕它们的编排代码。这使得自动化线束优化——人工智能系统对线束进行迭代和评估引导的改进——既是改进人工智能系统的重要途径,也是对人工智能系统本身的要求很高的能力。然而,社区缺乏一个通用的协议来衡量前沿法学硕士在这项任务上的表现。我们引入了 HarnessOpt-Bench,,这是在昂贵且随机的评估下进行端到端线束优化的基准。与编码工具, 配对的LLM 优化器, 接收目标代理的 种子工具, 分级评估反馈, 和固定目标评估预算。它编辑工具并提名最终候选,,该候选者通过其相对于保留测试分区上种子的归一化增益进行评分,而该测试分区在整个搜索过程中仍然无法访问。受信任的执行环境强制执行评估边界, 计量目标代理资源使用, 并保留候选版本以供审核。我们在 111 次评分运行中评估了 5 个前沿 LLM 作为优化器,既在共享编码工具下,又在其本机工具下跨 4 个下游任务,。实验结果表明,优化器模型比它们通过, 执行的编码工具分离的更多,本机工具并不总是优越,,并且不同任务和种子机制的增益差异很大。这些结果将线束优化确立为一种可测量的、具有区分性的能力,具有很大的改进空间。
As LLMs are increasingly deployed within agentic systems, their capabilities depend not only on the model weights but also on the harness: the prompts, tools, control flow, memory, and orchestration code surrounding them. This makes automated harness optimization -- the iterative and evaluation-guided improvement of a harness by an AI system -- both an important route to improving AI systems and a demanding capability for AI systems themselves. Yet the community lacks a common protocol for measuring how well frontier LLMs perform at this task. We introduce HarnessOpt-Bench, a benchmark for end-to-end harness optimization under expensive and stochastic evaluation. An optimizer, an LLM paired with a coding harness, receives a target agent的 seed harness, graded evaluation feedback, and a fixed target-evaluation budget. It edits the harness and nominates a final candidate, which is scored by its normalized gain over the seed on a held-out test partition that remains inaccessible throughout search. A trusted execution environment enforces the evaluation boundary, meters target-agent resource use, and preserves candidate versions for audit. We evaluate 5 frontier LLMs as optimizers both under a shared coding harness and under their native harnesses across 4 downstream tasks, over 111 scored runs. Experiment results show that optimizer models separate more than the coding harnesses they act through, native harnesses are not consistently superior, and gains vary substantially across tasks and seed regimes. These results establish harness optimization as a measurable and discriminative capability with large space for improvement.
科目: 人工智能 (cs.AI); 计算和语言 (cs.CL); 机器学习 (cs.LG)
Subjects: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)