arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

当可解码性不再足够:语言模型中的逻辑有效性表征、行为分离与因果测试

When Decodability Is Not Enough: Logical Validity Representations, Behavioral Dissociation, and Causal Tests in Language Models

Smitha Muthya Sudheendra, Jaideep Srivastava

arXiv 2609.02438首次发表:更新:

发表机构

University of Minnesota, Twin Cities(明尼苏达大学双城分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究针对5个开放权重Transformer模型,揭示语言模型中逻辑有效性的表征、行为表达与因果运用存在分离,其有效性信息可从隐藏状态强解码却未必在输出中体现。

AI 中文摘要

大型语言模型看似具备逻辑推理能力,但仅靠正确或错误的答案几乎无法说明模型的内部表征。本研究针对5个开放权重Transformer模型,采用在推理族、语义域、模板和难度水平上存在差异的匹配有效-无效前提-主张对开展逻辑验证研究。尽管模型行为表现接近随机水平,逻辑有效性却往往能从隐藏状态中近乎完美地解码,且在保留的模板、域和推理族下仍具有强可解码性;在正确性条件评估定义明确的条件下,即使是行为不正确的示例,有效性也保持高可解码性。同时,穷尽式留一测试揭示了该泛化存在明显局限,且与随机对照相比,沿探针衍生的有效性方向干预仅产生微弱、非特异性影响。研究结果表明,表征有效性、在行为中表达有效性及因果运用有效性三者是不同的,与有效性相关的信息可从模型隐藏状态中强解码,却不一定能可靠地在其输出中表达。

英文摘要

Large language models can look capable of logical reasoning, but correct or incorrect answers alone tell us little about what the model represents internally. We study logical verification in five open-weight transformer models using matched valid--invalid premise--claim pairs that vary across inference families, semantic domains, templates, and difficulty levels. Despite near-chance behavioral performance, logical validity is often almost perfectly decodable from hidden states and remains strongly decodable under held-out templates, domains, and inference families. Validity also remains highly decodable on behaviorally incorrect examples in the conditions where correctness-conditioned evaluation is well defined. At the same time, exhaustive leave-one-out tests reveal clear limits to this generalization, and interventions along probe-derived validity directions have only weak, nonspecific effects compared with random controls. Our results suggest that representing validity, expressing it in behavior, and using it causally are distinct. Validity related information can be strongly decodable from a model's hidden states without being reliably expressed in its output.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑