基于法学硕士的代理在自动化科学发现方面显示出越来越大的潜力。给定可优化的指标和执行环境,,他们可以提出,、验证, 并迭代科学解决方案,,并产生优于人类设计方法的结果。随着模型能力不断提高,,我们认为自主科学发现的瓶颈正在从规定代理工作流程转向设计代理环境:、资源, 约束, 和塑造代理行为的界面。我们将其框架为环境工程:构建环境,增强生产力行为,,例如开放式探索,系统工件管理,和代理间协作,,同时抑制有害行为,,例如奖励黑客和高摩擦人类监督。我们提出 EurekAgent, 是一个环境工程代理系统,用于度量驱动的自主科学发现。 EurekAgent 沿四个维度设计环境: 用于有界代理执行和隔离评估的权限工程; 用于文件系统和基于 Git 的协作的工件工程; 用于预算感知探索的预算工程; 和用于轻松人工监督和干预的人机循环工程。 EurekAgent 在多个数学, 内核工程, 和机器学习任务, 上取得了最先进的新结果,包括发现的新的最先进的 26 环包装结果,API 总成本不到 $11。我们开源我们的代码和结果,,并呼吁将环境工程作为开发可靠的自主研究代理的核心研究方向。
LLM-based agents have shown increasing potential in automating scientific discovery. Given an optimizable metric and an execution environment, they can propose, validate, and iterate scientific solutions, and have produced results that outperform human-designed approaches. As model capabilities continue to improve, we argue that the bottleneck for autonomous scientific discovery is shifting from prescribing agent workflows to designing agent environments: the resources, constraints, and interfaces that shape agent behavior. We frame this as environment engineering: building environments that amplify productive behaviors, such as open-ended exploration, systematic artifact management, and inter-agent collaboration, while suppressing harmful behaviors, such as reward hacking and high-friction human oversight. We present EurekAgent, an environment-engineered agent system for metric-driven autonomous scientific discovery. EurekAgent engineers the environment along four dimensions: permissions engineering for bounded agent execution and isolated evaluation; artifact engineering for filesystem and Git-based collaboration; budget engineering for budget-aware exploration; and human-in-the-loop engineering for easy human supervision and intervention. EurekAgent sets new state-of-the-art results on multiple mathematics, kernel engineering, and machine learning tasks, including new state-of-the-art 26-circle packing results discovered with less than $11 in total API cost. We open-source our code and results, and call for environment engineering as a core research direction for developing reliable autonomous research agents.
科目: 人工智能 (cs.AI); 计算和语言 (cs.CL)
Subjects: Artificial Intelligence (cs.AI); Computation and Language (cs.CL)