大型语言模型 (LLMs) 越来越多地在分子属性基准, 上进行评估,但准确性无法区分预测属性的模型和检索已发布数字的模型。我们审核了 12 个回归基准上的 22 个前沿模型进行逐字检索,发现它在五个数据集上广泛但相对基准特定:,超过 $50\%$ 的 LLM 显示逐字检索,,而在其余数据集上,它仅出现在孤立的单元格中。我们在两个推理水平上进行实验,发现推理改变了检索。对相同分子和相同提示, 进行的相同实验, 在较高推理水平上比在最低推理水平上更常被标记为$89\%$。最后,,我们测试了一种在污染最严重的情况下中断检索的方法,,并发现在某些情况下最强的模型仍然可以识别转换后的 SMILES 字符串和原始标签的组合。此外,, 抑制检索使不同模型的预测误差相对而言更加接近,,而它们对逐字检索的不同使用则将它们分开。这表明法学硕士的一般预测能力不仅仅取决于记忆值的数量。这项工作概述了使用法学硕士的分子回归基准中逐字检索的数量和深度。
Large language models (LLMs) are increasingly evaluated on molecular property benchmarks, but accuracy cannot distinguish a model that predicts a property from one that retrieves a published number. We audit 22 frontier models on 12 regression benchmarks for verbatim retrieval and find that it is widespread but relatively benchmark-specific: on five datasets more than $50\%$ of the LLMs show verbatim retrieval, while on the remaining datasets it appears only in isolated cells. We run our experiments at two reasoning levels and find that reasoning changes retrieval. The same experiments, on the same molecules and with the same prompt, are flagged $89\%$ more often at the higher reasoning level than at the lowest one. Finally, we test a way to interrupt retrieval in our most contaminated cases, and find that the strongest models in some cases still recognise a combination of transformed SMILES strings and original labels. Furthermore, suppressing retrieval moves the prediction errors of the different models closer together in relative terms, while their differing use of verbatim retrieval spreads them apart. This indicates that the general predictive capability of an LLM is not determined solely by the amount of memorised values. This work provides an overview of the amount and depth of verbatim retrieval in molecular regression benchmarks using LLMs.
科目:人工智能(cs.AI)
Subjects: Artificial Intelligence (cs.AI)