基于 LLM 的代理正在迅速被用于科学数据分析, 自动化任务,这些任务曾经受到人力时间和专业知识的限制。这种能力通常被认为是加速发现,,但它也加速了熟悉的故障模式,,快速生成看似合理的,,无限可修改的分析很容易生成,,有效地将假设空间转变为由选择性选择的分析支持的候选主张,,针对可发布的积极结果进行了优化。与软件, 不同,科学知识不是通过代码的迭代积累和事后统计支持来验证的。对单个数据集的流畅解释或显着结果并不是验证。因为缺失的证据是负空间, 那些可能会证伪该声明的实验和分析从未进行过或从未发表过。因此,我们建议在证伪第一标准:代理协助下产生的非实验性主张进行评估,代理不应主要用于制作最引人注目的叙述,,而应积极寻找声明可能失败的方式。

LLM-based agents are rapidly being adopted for scientific data analysis, automating tasks once limited by human time and expertise. This capability is often framed as an acceleration of discovery, but it also accelerates a familiar failure mode, the rapid production of plausible, endlessly revisable analyses that are easy to generate, effectively turning hypothesis space into candidate claims supported by selectively chosen analyses, optimized for publishable positives. Unlike software, scientific knowledge is not validated by the iterative accumulation of code and post hoc statistical support. A fluent explanation or a significant result on a single dataset is not verification. Because the missing evidence is a negative space, experiments and analyses that would have falsified the claim were never run or never published. We therefore propose that non-experimental claims produced with agentic assistance be evaluated under a falsification-first standard: agents should not be used primarily to craft the most compelling narrative, but to actively search for the ways in which the claim can fail.

科目:人工智能(cs.AI)

Subjects: Artificial Intelligence (cs.AI)