AI 代理是工具, 合著者, 还是研究人员? 我们提出了一项量化案例研究 ($N=1$): 一位监督 AI 编码代理的物理学家 (Claude Code, Sonnet 和 Opus 模型) 经过 12 个工作日和 57 个会话来构建 CLAX-PT, 一个可微分的单环扰动理论模块贾克斯。我们按照干预级别记录并分类了 15 起监督事件。该代理通过迭代预言机测试自主解决了十个问题。另外两个由物理学家的 领域知识提供。它不能的三个 - 都逃避了预言机检测 - 具有一个共同的属性:,代理将症状减轻视为根本原因解决。它花费了 57 个会话中的 33 个来调整代码架构中的系数,该代码架构无法代表目标物理,,并且即使提示重新考虑; 也无法重新评估其 CLASS-PT 分支选择,只有注入的物理概念(各向异性 BAO 阻尼) 触发了重新设计。另外, 代理进行了校准修正,通过了所有预言机测试,但与理论中的任何数量, 预测任何其他宇宙学的错误值无关。在同一个会话中,捏造因素被捕获并被替换。事实证明,三种监督实践对于捕捉预言机测试遗漏的内容至关重要: 在基准校准之外的不同参数点进行测试; 共享变更日志,这些变更日志显示了跨会话的停滞探索; 以及针对非物理数字补丁的明确规则。在这种情况下,监督设计,不是模型能力,决定了代理'的输出是否值得信赖。缩小差距需要代理提出架构替代方案,而不是在给定结构内进行优化,,并区分预测充分性和解释正确性——此处未展示的能力, 显然不能通过单独扩展来解决。 [A已桥接。]
Are AI agents tools, co-authors, or researchers? We present a quantified case study ($N=1$): a physicist supervising an AI coding agent (Claude Code, Sonnet and Opus models) over 12 work days and 57 sessions to build CLAX-PT, a differentiable one-loop perturbation theory module in JAX. We documented and classified 15 supervision events by intervention level. The agent resolved ten autonomously by iterating against oracle tests. Two more by the physicist的 domain knowledge. The three it could not -- all evaded oracle detection -- share a common property: the agent treated symptom reduction as root-cause resolution. It spent 33 of the 57 sessions adjusting coefficients within a code architecture that could not represent the target physics, and could not re-evaluate its CLASS-PT branch choice even when prompted to reconsider; only an injected physics concept (anisotropic BAO damping) triggered the redesign. Separately, the agent committed a calibrated correction that passed all oracle tests but corresponded to no quantity in the theory, predicting wrong values at any other cosmology. The fudge factor was caught and replaced within the same session. Three supervision practices proved critical for catching what oracle tests missed: testing at diverse parameter points beyond the fiducial calibration; shared changelogs that surfaced stalled exploration across sessions; and an explicit rule against unphysical numerical patches. In this case, supervision design, not model capability, determined whether the agent的 output was trustworthy. Closing the gap would require agents that propose architectural alternatives rather than optimize within a given structure, and distinguish predictive adequacy from explanatory correctness -- capabilities not exhibited here, not obviously addressed by scaling alone. [Abridged.]
科目: 人工智能 (cs.AI); 宇宙学和非银河天体物理学 (astro-ph.CO); 人机交互 (cs.HC); 软件工程 (cs.SE)
Subjects: Artificial Intelligence (cs.AI); Cosmology and Nongalactic Astrophysics (astro-ph.CO); Human-Computer Interaction (cs.HC); Software Engineering (cs.SE)