arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.18570cs.CLcs.AI

出于什么原因?解释模型对因果关系和对立关系的编码

For What Reason? Interpreting Models' Encoding of Causation and Antithesis

发表机构马萨诸塞大学阿默斯特分校 · 南佛罗里达大学
查看机构详情
  • University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校)
  • University of South Florida(南佛罗里达大学)

机构由 AI 辅助整理,请以论文原文为准。

Abhidip Bhattacharyya, Shira Wein

首次发表
浏览论文内容

中文总结 AI 辅助

研究指令微调的Transformer模型对英语话语关系的编码,聚焦因果和对立关系。通过下一个token预测任务及可解释性技术,发现早期层在序列中间做预测,中层接近末尾确定决策,部分层有答案偏好,揭示话语推理的不对称表示。

中文摘要 AI 辅助

话语关系为文档提供结构,对语言理解至关重要,并影响语言模型的性能和伦理。本文研究指令微调的Transformer模型(LLaMA和Mistral)如何对英语中的话语关系进行编码,尤其关注因果和对立的对比关系。将任务构建为下一个token预测任务并应用一系列可解释性技术测试模型内部。结果表明某些早期层在序列中间token处做出预测决策,一些中层在接近最后token时确定决策,其余层主要传播早期决策。还观察到一些层对一个答案有偏好,表明基于话语推理的不对称表示。

英文摘要

Discourse relations provide document structure, critical to language understanding and enabling language model performance and ethicality. In this work, we investigate how instruction-tuned Transformer models (LLaMA and Mistral) encode discourse relations in English, with a particular focus on the contrasting relations of causation and antithesis. Framing the task as a next-token prediction task and applying a suite of interpretability techniques to test model internals, our findings show that certain early layers make predictive decisions at mid-sequence tokens, while some mid-level layers finalize their decisions closer to the last token. Most of the remaining layers primarily propagate earlier decisions rather than actively influencing them. Additionally, we observe that some layers exhibit a preference for one answer over alternatives, suggesting asymmetric representation of discourse-based reasoning.\footnote{Our code is available at https://github.com/abhidipbhattacharyya/causation_vs_antithesis}

↑