arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

大型语言模型(LLMs)知道约束但不使用它:语用约束推理中的激活瓶颈

LLMs Know the Constraint But Do Not Use It: Activation Bottlenecks in Pragmatic Constraint Reasoning

Yubo Li, Ramayya Krishnan, Rema Padman

arXiv 2608.12321首次发表:更新:

发表机构

Carnegie Mellon University(卡内基梅隆大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究揭示LLMs虽编码了语用约束知识,但存在激活传递瓶颈,导致约束推理失败,激活修补可部分修复,且缓解干预未解决该传递问题。

AI 中文摘要

当显著的表面线索与隐含的可行性约束冲突时,大型语言模型(LLMs)常出现失败,但聚合准确率混淆了真正的约束推理与保守默认行为。我们将这种区别形式化为条件约束激活:约束在内部被编码(知识),在存在和不存在约束的提示中对称(对称性),但仅有时会被传递到决策(传递),并可通过补充激活修复(修复)。对14个模型的四元诊断显示两种失败模式;对两个开放权重的探测解码约束准确率超88%,但激活修补可修复其中一个(+6.4 nat),另一个则不行(-0.07)。在缓解前沿,无提示干预能达到修复角:所有都通过单一中介途径(前提提及)放大保守偏差。隐藏约束失败是传递问题,而非知识问题。

英文摘要

When a salient surface cue competes with an implicit feasibility constraint, LLMs often fail -- but aggregate accuracy conflates genuine constraint inference with conservative defaulting. We formalize the distinction as conditional constraint activation: the constraint is internally encoded (Knowledge) symmetrically across constraint-present and -absent prompts (Symmetry), yet only sometimes routed into the decision (Routing) and repairable by a donor activation (Repair). A quartet diagnostic over 14 models reveals two failure modes; probes on two open weights decode the constraint above $88\%$, yet activation patching repairs one ($+6.4$ nats) and not the other ($-0.07$). On a mitigation frontier, no prompted intervention reaches the repair corner: all inflate conservative bias through a single mediation pathway -- prerequisite mention. Hidden-constraint failure is a routing problem, not a knowledge problem.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑