个性化改变了模型对用户所说的内容; 我们表明,它还可以改变用于证明响应合理性的推理轨迹。现代法学硕士通过存储用户属性, 偏好, 和先前的上下文, 然后将此信息注入到未来的提示中来个性化交互。我们研究这种记忆是否会重塑对不存在单一基本事实答案的开放式问题的推理。为了量化这种效果,,我们引入了 DRIFTLENS, 一个与事实无关的框架,它将每个表达的推理步骤映射到一个值类别,并测量问题'的无记忆轨迹与其在注入用户属性记忆下的轨迹之间的差异。我们首先验证 DRIFTLENS 能够区分内容无关的语用噪声和实质性推理变化。在四个法学硕士和 10 个用户属性类别, 中,包括年龄, 职业, 和残疾, 用户属性记忆会导致每个模型的 实用噪声层, 之上出现中到大的推理漂移,即使最终答案仍然流畅, 主题, 且合理。然后,我们评估基于 GRPO 和 DPO 的训练后方法以减少漂移。两者都减少了漂移,,但都没有统一主导; 对下游能力, 有用性, 的影响, 和指令遵循是模型和奖励相关的。这些结果表明,记忆引起的推理漂移是个性化语言模型的一种可测量且仅部分减轻的故障模式。

Personalization changes what a model says to a user; we show that it can also change the reasoning trajectory used to justify the response. Modern LLMs personalize interactions by storing user attributes, preferences, and prior context, then injecting this information into future prompts. We study whether such memory reshapes reasoning on open-ended questions where no single ground-truth answer exists. To quantify this effect, we introduce DRIFTLENS, a ground-truth-free framework that maps each expressed reasoning step to a value category and measures divergence between a question的 no-memory trajectory and its trajectory under injected user-attribute memory. We first validate that DRIFTLENS distinguishes content-free pragmatic noise from substantive reasoning changes. Across four LLMs and 10 user-attribute categories, including age, occupation, and disability, user-attribute memory induces medium-to-large reasoning drift above each model的 pragmatic-noise floor, even when final answers remain fluent, on-topic, and plausible. We then evaluate GRPO- and DPO-based post-training methods for reducing drift. Both reduce drift, but neither uniformly dominates; effects on downstream capability, helpfulness, and instruction following are model-and reward-dependent. These results suggest that memory-induced reasoning drift is a measurable and only partly mitigated failure mode of personalized language models.

科目:人工智能(cs.AI)

Subjects: Artificial Intelligence (cs.AI)