为了测试正确的逻辑判断如何响应学习的上下文,,我们在保持模型固定的同时,在精确标记的三段论推理基准之前添加一个软前缀。软前缀是不透明的连续向量,,因此我们通过它们在逻辑形式和接口的受控变化中引起的行为来表征它们。通过研究哪些前缀成功以及它们的效果如何泛化,,我们描述了习得的上下文压力如何超越正确的判断并暴露模型的 逻辑稳定性的限制。在 Qwen3.6-35B-A3B MoE, Qwen3-8B, 和 Gemma 4 31B, 中,学习到的前缀会重定向许多正确答案,并在看不见的表单和界面更改中保持有效。在使用 Qwen3.6 MoE 和 Gemma, 进行的重复测试中,它们在所有 16 个模型方向分割比较中均优于配对随机对照 37 至 99 个百分点。 Qwen3.6 MoE 的翻转率在措辞和提示更改, 之间保持在 72% 到 90% 之间,而 Gemma 有效性前缀保留 54% 到 56% 的翻转率,而匹配的随机前缀的翻转率低于 1%。诊断测试表明,主导效应是对一种答案含义的广泛偏好,而不是固定符号强制或在任务之间可靠转移的逻辑操作。这种偏差的形式因模型而异。在两个 Qwen 模型, 中,简单评分模型通常会预测哪些判断会发生翻转,但不会预测其利润率会移动多远,,而 Gemma的总体响应更接近于相同模型。这些结果表明,成功的软前缀的主要行为效应是广泛的答案偏好,,而其余的响应则揭示了逻辑稳定性方面特定于模型的实质性差异。
To test how correct logical judgments respond to learned context, we prepend a soft prefix to an exactly labeled syllogistic reasoning benchmark while keeping the model fixed. Soft prefixes are opaque continuous vectors, so we characterize them through the behavior they induce across controlled variations in logical form and interface. By studying which prefixes succeed and how their effects generalize, we characterize how learned contextual pressure can override correct judgments and expose limits in a model的 logical stability. Across Qwen3.6-35B-A3B MoE, Qwen3-8B, and Gemma 4 31B, learned prefixes redirect many correct answers and remain effective across unseen forms and interface changes. In repeated tests with Qwen3.6 MoE and Gemma, they outperform paired random controls in all 16 model--direction--split comparisons by 37 to 99 percentage points. Qwen3.6 MoE flip rates remain between 72% and 90% across wording and prompt changes, while Gemma validity prefixes retain 54% to 56% flip compared with less than 1% for matched random prefixes. Diagnostic tests show that the dominant effect is a broad preference for one answer meaning rather than fixed-symbol forcing or a logical operation that transfers reliably between tasks. The form of this bias differs across models. In both Qwen models, simple score models often predict which judgments will flip but not how far their margins will move, whereas Gemma的 overall response is more closely approximated by the same models. These results show that the dominant behavioral effect of successful soft prefixes is a broad answer preference, while the remaining response reveals substantial model-specific differences in logical stability.
科目: 人工智能 (cs.AI); 计算和语言 (cs.CL)
Subjects: Artificial Intelligence (cs.AI); Computation and Language (cs.CL)