arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.18228cs.AIcs.CL

压力下的逻辑判断:用学习到的软前缀诊断三段论稳定性

Logical Judgments Under Pressure: Diagnosing Syllogistic Stability with Learned Soft Prefixes

Brian K Chen

首次发表
浏览论文内容

中文总结 AI 辅助

研究在三段论推理基准前加软前缀,探究学习到的上下文压力对模型逻辑判断的影响,发现软前缀能改变正确答案,在多模型测试中效果优于随机对照,主要影响是答案偏好,不同模型有显著差异。

中文摘要 AI 辅助

为测试正确的逻辑判断如何响应学习到的上下文,我们在一个精确标注的三段论推理基准前添加软前缀,同时保持模型不变。软前缀是不透明的连续向量,通过它们在逻辑形式和接口的受控变化中引发的行为来表征。通过研究哪些前缀成功及其效果如何泛化,我们刻画了学习到的上下文压力如何推翻正确判断并揭示模型逻辑稳定性的局限性。在多个模型上,学习到的前缀改变了许多正确答案,在未见过的形式和接口变化中依然有效,且在多次测试中优于随机对照。诊断测试表明,主要影响是对一种答案含义的广泛偏好,不同模型的这种偏差形式不同。这些结果表明,成功的软前缀的主要行为影响是广泛的答案偏好,同时其余响应揭示了逻辑稳定性方面模型特定的显著差异。

英文摘要

To test how correct logical judgments respond to learned context, we prepend a soft prefix to an exactly labeled syllogistic reasoning benchmark while keeping the model fixed. Soft prefixes are opaque continuous vectors, so we characterize them through the behavior they induce across controlled variations in logical form and interface. By studying which prefixes succeed and how their effects generalize, we characterize how learned contextual pressure can override correct judgments and expose limits in a model's logical stability. Across Qwen3.6-35B-A3B MoE, Qwen3-8B, and Gemma 4 31B, learned prefixes redirect many correct answers and remain effective across unseen forms and interface changes. In repeated tests with Qwen3.6 MoE and Gemma, they outperform paired random controls in all 16 model--direction--split comparisons by 37 to 99 percentage points. Qwen3.6 MoE flip rates remain between 72% and 90% across wording and prompt changes, while Gemma validity prefixes retain 54% to 56% flip compared with less than 1% for matched random prefixes. Diagnostic tests show that the dominant effect is a broad preference for one answer meaning rather than fixed-symbol forcing or a logical operation that transfers reliably between tasks. The form of this bias differs across models. In both Qwen models, simple score models often predict which judgments will flip but not how far their margins will move, whereas Gemma's overall response is more closely approximated by the same models. These results show that the dominant behavioral effect of successful soft prefixes is a broad answer preference, while the remaining response reveals substantial model-specific differences in logical stability.

发表机构

  • National University of Singapore(新加坡国立大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑