随着 LLM 采用变得更加广泛,,人们越来越有兴趣检测 LLM 生成的内容,,例如通过 LLM 检测工具和基于语言模式的启发式方法。检测器作为一种干预措施,不仅引导检测到的属性本身,,还引导下游指标,例如 LLM 使用情况和输出质量。在这项工作,中,我们展示了不完美的LLM检测器如何通过扭曲用户在工作流程中使用LLM的激励方式,对这些下游指标,产生违反直觉的影响。我们开发了一个程式化模型,该模型捕获用户如何策略性地选择使用 LLM 的程度以及如何对内容进行后处理以减少检测到的属性。使用这个模型,,我们表明 LLM 检测可以违反直觉地导致人们增加他们的 LLM 使用率。此外,即使减少检测到的属性也可以提高输出质量,我们发现引入LLM检测器可能会导致用户产生较低质量的输出。相比之下,,我们表明检测器会为检测到的属性, 产生干净的"rise-then-fall" 模式,我们根据经验在 arXiv 摘要上重现词频。总之,我们的工作说明了LLM检测如何扭曲LLM的使用和输出质量,当LLM检测器作为对这些下游指标的干预时发现故障模式。
As LLM adoption becomes more widespread, there is a growing interest in detecting LLM-generated content, for example through LLM detection tools and through heuristics based on language patterns. Detectors operate as an intervention that steers not only the detected attribute itself, but also downstream metrics such as LLM usage and output quality. In this work, we demonstrate how imperfect LLM detectors lead to counterintuitive impacts on these downstream metrics, by distorting how users are incentivized to use LLMs in their workflow. We develop a stylized model which captures how users strategically choose how much to use the LLM and how to post-process content to reduce the detected attribute. Using this model, we show that LLM detection can counterintuitively lead humans to increase their LLM usage. Moreover, even when reducing the detected attribute improves output quality, we find that introducing an LLM detector can lead users to produce lower quality outputs. In contrast, we show that detectors result in a clean "rise-then-fall" pattern for the detected attribute, which we empirically reproduce for word frequencies on arXiv abstracts. Altogether, our work illustrates how LLM detection can distort LLM usage and output quality, uncovering failure modes when LLM detectors operate as an intervention on these downstream metrics.
科目: 人工智能 (cs.AI); 计算机科学与博弈论 (cs.GT)
Subjects: Artificial Intelligence (cs.AI); Computer Science and Game Theory (cs.GT)