破坏预训练数据可能会给 LM 带来难以检测和缓解的有害行为。先前关于中毒预训练数据的工作很大程度上利用了维基百科,等已建立的数据源,这些数据源并不代表预训练语料库,典型的大规模和异质性,并且忽略了中毒数据和数据管理管道之间的交互。我们通过现有的网络规模内容注入机制:公共讨论接口证明,在超出此限制设置的情况下,对预训练数据的中毒攻击是可行的。另外, 为了衡量在网络爬行和数据管理之后是否包含恶意内容,,我们引入了 HalfLife, 一种新颖的分析方法,用于估计基于网络爬行的 LM 训练数据中包含的对抗性内容。我们使用 HalfLife 来探索通过开放讨论界面在网络规模上毒害预训练语料库的可行性。我们的分析证明了估计预训练数据,中是否包含毒物注入的重要性,并建立了第三方网页内容作为攻击语言模型预训练的可能向量。
Poisoning pretraining data can introduce harmful behaviors to LMs that are difficult to detect and mitigate. Prior work on poisoning pretraining data has largely exploited established data sources such as Wikipedia, which do not represent the large scale and heterogeneity typical of pretraining corpora, and has ignored the interaction between poisoned data and data curation pipelines. We demonstrate that poisoning attacks on pretraining data are feasible beyond this limited setting through an existing web-scale content injection mechanism: public discussion interfaces. Additionally, to measure whether malicious content is included after web crawling and data curation, we introduce HalfLife, a novel analysis for estimating adversarial content inclusion in web-crawl based LM training data. We use HalfLife to explore the feasibility of poisoning pretraining corpora at web scale through open discussion interfaces. Our analysis demonstrates the importance of estimating whether poison injections are included in pretraining data, and establishes third-party webpage content as a possible vector for attacking language model pretraining.
科目: 人工智能 (cs.AI); 计算和语言 (cs.CL)
Subjects: Artificial Intelligence (cs.AI); Computation and Language (cs.CL)