基准通过提供标准化和明确的绩效衡量标准,是评估和推进 LLM 和 MLLM 的基础。然而, 它们的构造是劳动密集型的并且难以重复使用, 引起了人们对可持续性和可扩展性的担忧。此外,现有基准测试在发布,后通常很快达到性能饱和,导致对最先进模型的区分不足。为了应对这些挑战,,我们引入了 Benchmark Agent,,这是一个专为基准构建而设计的完全自主的代理系统。我们的框架编排了从用户查询分析和子任务设计到数据注释和质量控制的完整基准构建管道,。为了评估基准代理,,我们实施它来生成 15 个代表性基准,,涵盖不同的评估场景,,包括文本理解, 多模式理解, 和特定领域推理。广泛的实验, 包括人工评估, LLM 作为法官评估, 和一致性检查, 证明 Benchmark Agent 可以在最少的人工参与下生成高质量的基准样本。更重要的是,通过持续评估,我们观察到了一些富有洞察力的发现,包括当前模型难以应对某些特定领域的推理任务。我们相信快速发展的基准可以为研究界做出重大贡献。预览和代码将在演示页面和代码存储库中公开提供。
Benchmarks are fundamental for evaluating and advancing LLMs and MLLMs by providing standardized and explicit measures of performance. However, their construction is labor-intensive and hard to reuse, raising concerns about sustainability and scalability. Moreover, existing benchmarks often quickly reach performance saturation after their release, resulting in insufficient discrimination among state-of-the-art models. To address these challenges, we introduce Benchmark Agent, a fully autonomous agentic system designed for benchmark building. Our framework orchestrates the complete benchmark construction pipeline, from user query analysis and subtask design to data annotation and quality control. To assess Benchmark Agent, we implement it to produce 15 representative benchmarks, spanning diverse evaluation scenarios, including text understanding, multimodal understanding, and domain-specific reasoning. Extensive experiments, including human evaluation, LLM-as-a-judge assessment, and consistency checks, demonstrate Benchmark Agent can generate high-quality benchmark samples with minimal human involvement. More importantly, through continual evaluation, we observe several insightful findings, including that current models struggle with certain domain-specific reasoning tasks. We believe that rapidly evolving benchmarks can contribute significantly to the research community. The preview and code will be publicly available at the demo page and code repository.
科目:人工智能(cs.AI)
Subjects: Artificial Intelligence (cs.AI)