arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.10268cs.CL

语言再生:信息局部性对重构影响的研究

Language Re-generation: An investigation into information locality effects on reconstruction

  • Utrecht University(乌得勒支大学)

机构由 AI 辅助整理,请以论文原文为准。

Amirhossein Mohammadi, Laurence E. Frank, Albert Gatt, Robert A. Bagheri

AI总结:

研究信息局部性对语言模型恢复自然语言的影响,通过微调预训练的GPT-2模型从三种扰动类型重构自然英语,发现恢复结构有局部性偏好,揭示架构偏差,且恢复难度与可学习性难度相关,表明信息局部性是共同约束。

AI中文摘要:

信息局部性,即句法相关词汇倾向于相邻出现,影响人类语言处理和语言模型学习。此前研究了语言模型能否习得不可能的语言,但不清楚它们能否从这类输入中恢复自然语言及其揭示的归纳偏差。我们通过用一个重构框架补充基于可学习性的方法来解决此问题:微调在不可能的语言上预训练的GPT-2模型,从三种扰动类型中重构自然英语。结果表明,恢复的结构比原文依赖长度短,反映了无约束语言模型生成中的局部性偏好,揭示了仅靠可学习性实验未发现的架构偏差。恢复难度随局部性破坏程度增加。结构恢复(依存三元组F1)与表面恢复(完全匹配)分离,流畅性与全局洗牌下的忠实重构分离。句子长度进一步调节性能:局部结构保留时,长句子便于恢复,全局洗牌下则完全崩溃。最后,恢复难度在不同扰动类型中跟踪可学习性难度,表明信息局部性是两者共有的约束。

英文摘要:

Information locality, the tendency for syntactically related words to appear close together, shapes both human language processing and language model learning. While prior work has examined whether language models can acquire impossible languages, it remains unclear whether they can recover natural language from such input and what this reveals about their inductive biases. We address this by complementing learnability-based approaches with a reconstruction framework: fine-tuning GPT-2 models pre-trained on impossible languages to reconstruct natural English from three perturbation types. Our findings show that the recovered structures exhibit shorter dependency lengths than the original text, mirroring the locality preference observed in unconstrained language model generation and providing a quantitative signature of an architectural bias that learnability experiments alone do not reveal. Recovery difficulty increases with the degree of locality disruption. Structural recovery (dependency Triple F1) dissociates from surface recovery (Exact Match), while fluency dissociates from faithful reconstruction under global shuffling. Sentence length further modulates performance: longer sentences facilitate recovery when local structure is preserved but lead to complete collapse under global shuffling. Finally, recovery difficulty tracks learnability difficulty across perturbation types, suggesting that information locality is the shared constraint governing both.

↑