现代化遗留 Fortran 是一个体积: 的问题,转换是单独例行的,,但代码库可能是巨大的,,并且在许多计算科学中,工作根本就没有完成。我们提出了一个代理工作流程,以生产规模, 的方式开展这项工作,并且我们开始衡量这种委托可以达到多远。在这项工作,中,三个提示专业代理角色在版本控制的规范下运行,该规范是代理自己编写和修订的,,而人类持有少量的门。该安排通过从域,继承的精确验证预言机来保证安全,并且安全委托的边界恰好位于该预言机停止看到的地方。我们在案例研究,中应用了所提出的工作流程,将GAMESS (通用原子和分子电子结构系统),一个具有48年开发历史,的成熟量子化学包的二电子积分例程从固定形式的Fortran 77转换为自由形式的Fortran 2008。这项工作的范围是十二个源文件, 56,448行,和225用于计算电子斥力积分的子程序。这些代理在独立的工作树, 中作为三个 Claude 代码角色运行,并且工作跨越了四代 Claude 模型。因为 GAMESS 小组发布了一个标准测试套件,其打印能量被用户社区视为规范,,所以我们可以采用这些能量的逐位再现作为合并标准,,其中小数点后第 12 位的偏差被视为失败而不是漂移。所有 12 个源文件都通过了 51 项测试验证,其中包括 49 个标准 GAMESS 测试和两个附加计算,,并且在 612 个测试运行中,与化学相关的差异数量为零,,并且每个文件还通过了用于持续集成的 Jenkins 测试。
Modernizing legacy Fortran is a problem of volume: the transformations are individually routine, but the codebases can be enormous, and across much of computational science the work simply goes undone. We propose an agentic workflow that takes this work on at production scale, and we set out to measure how far such delegation can reach. In this work, three prompt-specialized agent roles operate under a version-controlled specification that the agents themselves authored and revised, while humans hold a small number of gates. The arrangement is kept safe by an exact verification oracle inherited from the domain, and the boundary of safe delegation lies exactly where that oracle stops seeing. We apply the proposed workflow in a case study, converting the two-electron-integral routines of GAMESS (General Atomic and Molecular Electronic Structure System), a mature quantum-chemistry package with a 48-year development history, from fixed-form Fortran 77 to free-form Fortran 2008. The scope of this work was twelve source files, 56,448 lines, and 225 subroutines for computing electron repulsion integrals. The agents ran as three Claude Code roles in isolated worktrees, and the work spanned four Claude model generations. Because the GAMESS group ships a standard test suite whose printed energies its user community treats as canonical, we could adopt bit-for-bit reproduction of those energies as the merge criterion, where a deviation in the twelfth decimal place counts as a failure rather than drift. All twelve source files pass a 51-test validation battery comprising the 49 standard GAMESS tests and two additional calculations, and across 612 test runs the number of chemistry-relevant differences is zero, and every file also passes the Jenkins tests that are used for continuous integration.
科目:人工智能(cs.AI)
Subjects: Artificial Intelligence (cs.AI)