代理系统正在跨领域快速发展,,但它们的评估仍然分散。大多数基准测试都依赖于固定, 以 LLM 为中心的工具,这些工具需要大量集成, 造成测试与生产不匹配, 并限制不同代理设计之间的公平比较。根本问题是缺乏开放的, 与代理无关的评估界面。我们提倡代理化代理评估 (AAA),,其中评估由法官代理执行,所有参与者通过用于任务管理的标准化协议: A2A 和用于工具访问的 MCP 进行交互。传统的基准测试定义了两个单独的接口,,一个用于基准测试,一个用于代理,,而 AAA 仅需要一个;,这产生了一个通用的, 统一框架,该框架将评估逻辑与代理实现分开,并实现可重复的, 可互操作, 和多代理评估。我们进一步介绍 AgentBeats 作为 AAA: 的具体实现,我们确定了五种实际操作模式,使标准化评估与现实世界对开放性, 隐私, 和可重复性的限制兼容。为了大规模评估我们的设计,,我们进行了两项研究:,这是一项为期五个月的公开竞赛,吸引了 12 个类别的 298 名评判代理以及来自独立参与者的 467 名主体代理,,表明 AAA 适用于各种不同的基准; 以及一项关于编码代理的案例研究,证实代理化评估保留了公共记录的保真度,同时揭示了以前缺失的面对面结果, 产生了有关代理设计的研究见解。结合社区规模的实地研究和受控编码案例研究,,我们验证了 AAA 提供了大规模异构场景的覆盖, 实用性, 和保真度。 , AAA 和 AgentBeats 共同提供了一条通向开放, 标准化, 和可重复代理评估的清晰路径。
Agent systems are advancing quickly across domains, but their evaluation remains fragmented. Most benchmarks rely on fixed, LLM-centric harnesses that require heavy integration, create test-production mismatch, and limit fair comparison across diverse agent designs. The root problem is the lack of an open, agent-agnostic assessment interface. We advocate Agentified Agent Assessment (AAA), where evaluation is performed by judge agents and all participants interact through standardized protocols: A2A for task management and MCP for tool access. Conventional benchmarking defines two separate interfaces, one for the benchmark and one for the agent, while AAA only needs one; this yields a generic, unified framework that separates assessment logic from agent implementation and enables reproducible, interoperable, and multi-agent evaluation. We further introduce AgentBeats as a concrete realization of AAA: we identify five practical operation modes that make standardized assessment compatible with real-world constraints on openness, privacy, and reproducibility. To evaluate our design at scale, we conduct two studies: a five-month open competition that drew 298 judge agents across 12 categories together with 467 subject agents from independent participants, showing that AAA applies across a heterogeneous range of benchmarks; and a case study on coding agents that confirms agentified evaluation preserves fidelity with the public record while surfacing previously missing head-to-head results, yielding research insights about agent design. Combining a community-scale field study and a controlled coding case study, we verify that AAA delivers coverage, practicality, and fidelity across heterogeneous scenarios at scale. Together, AAA and AgentBeats offer a clear path toward open, standardized, and reproducible agent assessment.
科目: 人工智能 (cs.AI); 机器学习 (cs.LG)
Subjects: Artificial Intelligence (cs.AI); Machine Learning (cs.LG)