长视野推理需要一个代理运行时,当证据支持其当前方法时,该运行时可以持续存在,并且当测量显示失败,隐藏约束,或错误指定的目标时,该运行时可以旋转。我们向 Argus, 展示了一个持久的, 自我演化运行时,其中 Manager, Planner, Engineer, 和 Reviewer 在持久项目状态上执行有界任务。 Argus 将稳定的用户意图与操作目标, 约束, 和验证标准, 分开,并承认记忆, 技能, 程序, 验证者, 路由决策, 和仅在角色拥有的审查和, 可用时, 任务本机验证之后拒绝路由。模型权重保持固定; 通过持久运行时状态和控制策略, 进行自我演化,并在操作员拥有的升级点之间自主执行。在七个 GPT-5.5 基准测试领域中,, Argus 在 SWE-Bench Pro 上获得了约 78%,而在 Direct Copilot 上获得了 59%,同时使用了 1.41 倍的总代币。在验证门控自我进化,之后,与启动wave,相比,成熟的SWE-Bench Wave使用的求解输入令牌减少了21%,每个任务的活动工作流程时间减少了15%,同时记录了34个验证者恢复和22个严格的审查循环救援。 Argus 在 AARRI-Bench 上也达到了 76.8%,在数学数据综合, 上与具有竞争力的 GPU 内核和语言模型训练结果相比有 28.0 分的差距。除了基准,,优化的 RWKV6 内核已合并到上游; 为期多天的数学活动保留了伪造的路线和有证据支持的前沿更新; 和六个纸质管道完成了 254 项任务,其中有 16 个阶段回滚。这些结果表明,固定权重,自进化工具可以修改,恢复,并积累经过验证的方法,同时为未来的监督和强化学习生成结构化轨迹。

Long-horizon reasoning requires an agentic runtime that can persist when evidence supports its current approach and pivot when measurements reveal failure, hidden constraints, or a misspecified objective. We present Argus, a persistent, self-evolving runtime in which Manager, Planner, Engineer, and Reviewer execute bounded missions over durable project state. Argus separates stable user intent from operational objectives, constraints, and verification criteria, and admits memories, skills, procedures, verifiers, routing decisions, and rejected routes only after role-owned review and, when available, task-native verification. Model weights remain fixed; self-evolution occurs through persistent runtime state and control policy, with autonomous execution between operator-owned escalation points. Across seven GPT-5.5 benchmark arenas, Argus achieves about 78% on SWE-Bench Pro versus 59% for Direct Copilot while using 1.41 times the aggregate tokens. After verification-gated self-evolution, mature SWE-Bench waves use 21% fewer solve-input tokens and 15% less active workflow time per task than startup waves, while recording 34 verifier recoveries and 22 strict review-loop rescues. Argus also reaches 76.8% on AARRI-Bench and a 28.0-point gap on mathematical data synthesis, with competitive GPU-kernel and language-model-training results. Beyond benchmarks, an optimized RWKV6 kernel was merged upstream; a multi-day mathematics campaign retained falsified routes and proof-backed frontier updates; and six paper pipelines completed 254 missions with 16 stage rollbacks. These results show that a fixed-weight, self-evolving harness can revise, recover, and accumulate verified approaches while producing structured trajectories for future supervised and reinforcement learning.

科目:人工智能(cs.AI)

Subjects: Artificial Intelligence (cs.AI)