科学发现本质上是一个创造性和不确定性的过程,,需要超出已知知识回忆的推理。虽然已经提出了许多基准来通过多跳检索来评估大型语言模型 (LLM) 在深度研究任务中的性能,,但它们对于真正的科学发现至关重要的创新推理能力在很大程度上仍未经过测试。我们引入了一个基准框架,用于评估科学发现和推理, 中的模型性能,从原始问题到经典的零假设检验。在我们的框架, 模型中,最初仅接收最近论文, 中的主题和研究问题,并逐步揭示技术细节。在信息披露,的每个阶段,模型的任务是生成解决研究问题,的假设,将其与原始论文的结论进行比较,并通过构成原子权利要求的自动语义相似性进行评估。这种对语义分歧与真实结论的渐进评估能够评估模型的在最少信息)下的创新性(到基于完整实验细节),的基础推理能力(,这对于使用法学硕士进行科学发现目的至关重要。我们的框架为系统地评估法学硕士,的科学推理和发现能力提供了基础,这对于推进下一代人工智能科学家/联合科学家系统的发展至关重要。具体来说,, 在这里,我们评估了 GPT-5, GPT-5.4, Gemini 2.5 pro, 和 Gemini 3.1 pro 预览版,涉及 45 篇论文,涵盖生物活性材料, 机械材料, 和纳米材料。我们发现 GPT-5.4 和 Gemini 3.1 pro 的表现优于上一代同类产品,正如预期的那样,,而且 GPT-5.4 特别是即使在最小的背景下也能保持 0.7 F1 分数与地面真实结论的一致性。
Scientific discovery is an inherently creative and uncertain process, requiring reasoning beyond the recall of known knowledge. While many benchmarks have been proposed to evaluate large language model (LLM) performance on deep research tasks via multi-hop retrieval, their innovative reasoning abilities essential for true scientific discovery remain largely untested. We introduce a benchmark framework for evaluating model performance in scientific discovery and reasoning, building up from a raw problem to the classical null hypothesis test. In our framework, models initially receive only the topic and research question from a recent paper, with technical details progressively revealed. At each stage of information disclosure, the model is tasked with generating hypotheses that address the research question, which is compared with the conclusions from the original paper and evaluated via automated semantic similarity of constituent atomic claims. This progressive evaluation of semantic divergence from ground-truth conclusions enables assessment of a model的 innovativeness (under minimal information) to grounded reasoning capabilities (under full experimental details), both critical for using LLMs for scientific discovery purposes. Our framework provides a foundation for systematically evaluating scientific reasoning and discovery capabilities in LLMs, crucial for advancing the development of next-generation AI scientist/co-scientist systems. Specifically, here we evaluate GPT-5, GPT-5.4, Gemini 2.5 pro, and Gemini 3.1 pro preview across 45 papers spanning bioactive materials, mechanical materials, and nanomaterials. We find that GPT-5.4 and Gemini 3.1 pro outperform their previous generation counterparts as expected, and GPT-5.4 in particular maintains 0.7 F1 score alignment with ground truth conclusions even under minimal context.
科目:人工智能(cs.AI)
Subjects: Artificial Intelligence (cs.AI)