虽然递归自我改进(RSI)的研究主要采用自动化模型训练管道,,但可靠的自主开发需要缺失的支柱:事后监控和审计,以了解模型学习的内容并确保安全对齐。机械可解释性工具对于弥合这一差距至关重要,,其中稀疏自动编码器(SAE) 通过隔离模型检查和转向的可解释特征作为基石。在本文,中,我们引入了 SAEScientist-Bench 来评估 AI 代理是否可以充当科学家,利用 SAE 工具进行自主机械发现。给定目标概念,,代理设计对比探针并导航 Gemma-2-9B-IT 中包含 131K+ 特征的 Gemma Scope 字典,以发现最佳特征,,根据锚定在 Neuronpedia 上的策划专家参考特征进行评估,跨激活等级, 对对比文本, 的概念选择性, 和因果转向。在 10 个代理配置和 20 个任务中,, 前沿代理展示了真正的发现能力,并领先于不同的评估维度,,但仍远远落后于专家基线, 在将目标概念与对比控制分开方面接近专家水平,同时在因果生成指导方面大幅落后。进一步的分析表明,尽管代理可以设计对比来排除虚假候选,,但他们经常会误解实验测量结果。这些结果将实验模型理解建立为闭环自主 AI R&D 的可测量能力。我们的代码可以在这个 https URL 上找到。
While research on recursive self-improvement (RSI) has predominantly automated model training pipelines, reliable autonomous development demands a missing pillar: post-hoc monitoring and auditing to understand what models learn and ensure safe alignment. Mechanistic interpretability tools are essential to bridge this gap, among which Sparse Autoencoders (SAEs) serve as a cornerstone by isolating interpretable features for model inspection and steering. In this paper, we introduce SAEScientist-Bench to evaluate whether AI agents can act as scientists utilizing SAE tools for autonomous mechanistic discovery. Given a target concept, an agent designs contrastive probes and navigates a Gemma Scope dictionary of 131K+ features in Gemma-2-9B-IT to discover the optimal feature, evaluated against curated expert reference features anchored on Neuronpedia across activation rank, concept selectivity on contrastive texts, and causal steering. Across 10 agent configurations and 20 tasks, frontier agents demonstrate genuine discovery capabilities and lead different evaluation dimensions, but remain well behind the expert baseline, approaching expert levels on separating target concepts from contrastive controls while lagging substantially in causal generation steering. Further analysis reveals that although agents can design contrasts to rule out spurious candidates, they frequently misinterpret experimental measurements. These results establish experimental model understanding as a measurable capability for closed-loop autonomous AI R&D. Our code is available at this https URL.
科目: 人工智能 (cs.AI); 计算和语言 (cs.CL); 机器学习 (cs.LG)
Subjects: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)