arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.20647cs.CL

用于依存关系的定向上下文表示:为何跨方向配对会失效

Directional Contextual Representations for Dependency Relations: Why Cross-Direction Pairing Fails

发表机构人工智能与数据科学研究院(IAD) · 布法罗大学
查看机构详情
  • Institute for Artificial Intelligence and Data Science (IAD)(人工智能与数据科学研究院(IAD))
  • University at Buffalo(布法罗大学)

机构由 AI 辅助整理,请以论文原文为准。

Sai Krishna Arthanari, JaeHyeong Chang, Chengzhe Sun, Siwei Lyu

首次发表
浏览论文内容

中文总结 AI 辅助

本研究针对双向LSTM拆分为正向、反向状态的方法,发现跨方向配对表现逊于同方向配对且差距随标记距离增大,通过冻结主干等诊断揭示方向信息的存储与传播特性,解释了跨方向配对失效的部分机制。

中文摘要 AI 辅助

将双向LSTM的上下文表示拆分为仅正向的$F_i$(严格依赖于标记1..i)和仅反向的$B_i$(严格依赖于标记i..n),在依存关系类型分类任务上的表现优于单独使用其中任意一种,也优于融合自注意力的表示。但该思路的一个特定自然扩展——将一个标记的正向状态与候选标记的反向状态配对(即“跨方向”配对,$F_i$对$B_j$),其表现始终逊于同方向配对,且这种性能损失会随标记距离增大而加剧,两项结果均通过配对自助法验证具有统计显著性。本文采用冻结主干的方法诊断原因:通过代码检查确认所用为单层BiLSTM,从结构上排除了方向间的架构信息泄露;冻结主干仅训练新的头部时,93%的同方向与跨方向性能差距仍然存在,排除了训练协同适应是主要原因;线性回归显示$F_i$与$B_i$间存在部分表示冗余($R^2=0.324$,而打乱后的对照为0.028),线性探测显示$F_i$中存在对后续标记的部分预测编码(36.5%,多数基线为17.2%),这些都是真实存在的效应,但单独或结合都无法完全解释全部差距。扩展的冻结主干诊断(位置探测与距离衰减探测)显示,方向信息确实被存储但未精确定位,且仅能传播少数标记后就衰减至基线,这与距离增大时性能损失加剧的发现一致,并构成其机制基础。

英文摘要

Splitting a bidirectional LSTM's contextual representation into a forward-only $F_i$ (strictly a function of tokens $1..i$) and a backward-only $B_i$ (strictly a function of tokens $i..n$) beats either alone and beats a fused self-attention representation for dependency relation-type classification. But a specific, natural extension of this idea -- pairing a token's forward state against a \emph{candidate}'s backward state (``cross-direction'' pairing, $F_i$ vs.\ $B_j$) -- consistently \emph{underperforms} same-direction pairing, and the penalty \emph{grows}, not shrinks, with token distance, both paired-bootstrap significant. We diagnose why using a frozen-trunk methodology: architectural information leakage between directions is impossible by construction (a single-layer BiLSTM, verified by code inspection); 93\% of the same-vs-cross gap survives freezing the trunk and training only fresh heads, ruling out training-co-adaptation as the primary cause; linear regression shows partial representational redundancy between $F_i$ and $B_i$ ($R^2{=}0.324$ vs.\ $0.028$ for a shuffled control) and a linear probe shows partial anticipatory encoding of upcoming tokens in $F_i$ (36.5\% vs.\ 17.2\% majority baseline) -- real effects, but neither alone, nor combined, cleanly explains the full gap. Extended frozen-trunk diagnostics (a positional probe and a distance-decay probe) show directional information is genuinely stored but not exactly positioned, and propagates only a few tokens before decaying to baseline -- consistent with, and mechanistically underneath, the distance-growth finding.

↑