社会和行为科学的可重复性通常由独立研究人员评估,他们重新分析原始数据,以评估已发表的研究结果是否可以恢复。然而, 这种方法是资源密集型的并且难以扩展。在这里,我们展示了大型语言模型(LLMs)可以自动进行再现性评估。使用 N= 180 项已发表的研究以及来自行为和社会科学, 的预定义声明,我们将法学硕士生成的分析与原始结果进行比较。对于 11 项研究,,LLM 流程无法产生可行的效应大小估计。对于其余研究,,法学硕士在 80% 的病例, 中得出了与原始研究相同的定性结论,并在 24% 的研究中使用 Cohen的 d) 中的 +/-0.05 公差恢复了原始效应大小 (。在人类重新分析的子集中,,法学硕士在 95% 的研究, 中达到了与原始研究相同的定性结论,与人类再分析者相似(83%),,并且法学硕士在 40% 的研究中使用 +/-0.05 容差恢复了原始效应大小, 再次与人类再分析者大致相似(28%)。鉴于法学硕士, 目前的能力和局限性,研究结果表明法学硕士可以支持对实证结果的系统审计,而不是替代专家判断。因此,, 法学硕士可以作为一种可扩展的筛选工具,以提高实证研究的严谨性和可重复性。
Reproducibility in the social and behavioral sciences is typically evaluated by independent researchers who reanalyze the original data to assess whether the published findings can be recovered. However, such approaches are resource-intensive and difficult to scale. Here, we show that large language models (LLMs) can automate reproducibility assessments. Using N = 180 published studies with predefined claims from the behavioral and social sciences, we compare LLM-generated analyses with the original findings. For 11 studies, the LLM pipeline could not produce a viable effect size estimate. For the remaining studies, the LLM reached the same qualitative conclusion as the original study in 80% of cases, and recovered the original effect sizes (using a +/-0.05 tolerance in Cohen的 d) in 24% of studies. In a subset with human reanalyses, the LLM reached the same qualitative conclusion as the original study in 95% of studies, similar to human reanalysts (83%), and the LLM recovered the original effect sizes using a +/-0.05 tolerance in 40% of studies, again broadly similar to human reanalysts (28%). Given the current capabilities and limitations of LLMs, the findings show that LLMs can support systematic audits of empirical results rather than substitute expert judgment. As such, LLMs can serve as a scalable screening tool to improve the rigor and reproducibility in empirical research.
科目:人工智能(cs.AI)
Subjects: Artificial Intelligence (cs.AI)