大型语言模型 (LLMs) 越来越擅长数学推理,,但它们的不可靠性限制了它们在数学研究中的实用性。缓解措施是使用法学硕士以精益等语言生成正式证明。我们对该方法解决开放问题的能力进行了首次大规模评估。我们最有能力的代理自主解决了 353 个开放 Erdős 问题中的 9 个,每个问题的成本为几百美元, 证明了 44/492 OEIS 猜想,,并被部署在组合学, 优化, 图论, 代数几何, 和量子光学研究中。基本代理将基于 LLM 的生成与基于精益的验证交替使用,复制了 Erdős 的成功,但事实证明,在最困难的问题上成本更高。这些发现证明了人工智能辅助形式证明搜索的力量,并揭示了实现它的代理设计。

Large language models (LLMs) increasingly excel at mathematical reasoning, but their unreliability limits their utility in mathematics research. A mitigation is using LLMs to generate formal proofs in languages like Lean. We perform the first large-scale evaluation of this method的 ability to solve open problems. Our most capable agent autonomously resolved 9 of 353 open Erdős problems at the per-problem cost of a few hundred dollars, proved 44/492 OEIS conjectures, and is being deployed in combinatorics, optimization, graph theory, algebraic geometry, and quantum optics research. A basic agent alternating LLM-based generation with Lean-based verification replicated the Erdős successes but proved costlier on the hardest problems. These findings demonstrate the power of AI-aided formal proof search and shed light on the agent designs that enable it.

科目:人工智能(cs.AI)

Subjects: Artificial Intelligence (cs.AI)