人们越来越期望自主代理通过反馈来改进可执行策略,,但现有的评估经常将这一过程压缩为最终分数,或者将其与开放式软件工程进展相混淆。我们引入了自主策略演进,,这是一种受控评估设置,其中利用模型代理在固定交互预算下重复编辑可执行策略系统。我们在 EvoPolicyGym, 中实例化了此设置,这是一个从紧凑的交互式 RL 环境构建的基准,用于评估代理如何迭代地改进探索的策略。在 EvoPolicyGym 套件, GPT-5.5 上,在所有 16 个环境中实现了最强的总排名得分和前两名的性能。除了排行榜结果, EvoPolicyGym 还提供轨迹级诊断,区分代理如何分配预算, 将反馈转换为参数调整。这些分析表明,强大的自主策略演化不仅取决于孤立任务的胜利,,还取决于发现适合任务的机制并在有限反馈下完善策略。
Autonomous agents are increasingly expected to improve executable policies through feedback, yet existing evaluations often collapse this process into a final score or confound it with open-ended software-engineering progress. We introduce Autonomous Policy Evolution, a controlled evaluation setting in which a harness-model agent repeatedly edits an executable policy system under a fixed interaction budget. We instantiate this setting in EvoPolicyGym, a benchmark built from compact interactive RL environments that evaluates how agents iteratively improve explored policies. On the EvoPolicyGym suite, GPT-5.5 achieves the strongest aggregate rank score and top-two performance on all 16 environments. Beyond leaderboard results, EvoPolicyGym also provides trajectory-level diagnostics that distinguish how agents allocate budget, convert feedback into parametric tuning. These analyses show that strong autonomous policy evolution depends not only on isolated task wins, but on discovering task-appropriate mechanisms and refining policies under bounded feedback.
科目: 人工智能 (cs.AI); 计算和语言 (cs.CL)
Subjects: Artificial Intelligence (cs.AI); Computation and Language (cs.CL)