随着基础模型的进步和智能体支架变得越来越复杂,, 智能体在复杂的, 长期编码任务甚至自主实验执行方面表现出了卓越的熟练程度。尽管它们从研究助理演变为自主研究代理人,,这些系统在领域敏感性,研究道德,和细致入微的科学判断方面仍然表现出显着的局限性。因此,前沿特工仍然无法完全取代人类研究人员。为了弥补这一差距,,我们将 AARR (Act As a Real Researcher) 基准系列概念化。与主要评估宏观执行能力的现有基准不同,, AARR 重点关注代理是否能够模仿人类研究人员在精细研究场景中的专业性, 彻底性, 和细致入微的推理。在这项工作, 中,我们建议 AARRI-Bench (Act As a Real Research Intern), 作为本系列中的第一个基准。我们对前沿模型和代理系统, 进行了广泛的实验,结果表明,即使是性能最佳的配置 (Mini-SWE-Agent 和 Claude Opus 4.7) 也只能达到 68.3\% 的成功率, 经常忽略对真正的人类研究人员来说显而易见的微妙但关键的细节。我们的结果表明,开发类似研究人员的人工智能需要进一步探索研究行为,,而不仅仅是复杂的脚手架。我们的数据在此 https URL 发布。

As foundation models advance and agent scaffolding becomes increasingly sophisticated, agents have demonstrated remarkable proficiency in complex, long-horizon coding tasks and even autonomous experiment execution. Despite their evolution from research assistants into autonomous research agents, these systems still exhibit significant limitations in field sensitivity, research ethics, and nuanced scientific judgment. Consequently, frontier agents remain unable to fully replace human researchers. To bridge this gap, we conceptualize the AARR (Act As a Real Researcher) benchmark series. Unlike existing benchmarks that primarily assess macro-level execution capabilities, AARR focuses on whether agents can emulate the professionalism, thoroughness, and nuanced reasoning that characterize human researchers in granular research scenarios. In this work, we propose AARRI-Bench (Act As a Real Research Intern), the first benchmark in this series. We conduct extensive experiments across frontier models and agentic systems, revealing that even the best-performing configuration (Mini-SWE-Agent with Claude Opus 4.7) achieves only 68.3\% success rate, frequently overlooking subtle yet critical details that are obvious to real human researchers. Our results indicate that developing researcher-like AI requires further exploration of research behavior, rather than merely complex scaffolding. Our data is released at this https URL.

科目:人工智能(cs.AI)

Subjects: Artificial Intelligence (cs.AI)