arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

大型语言模型是否理解序列结构?一项关于推理与生成的受控研究

Do LLMs Understand Sequential Structure? A Controlled Study of Inference and Generation

Jerry Wang, Zhengxiang Wang, Ting Yu Liu, Hsin-Ling Hsu, Yi-Cheng Lai, Tengfei Ma

arXiv 2610.04977首次发表:更新:

发表机构

University of Illinois Urbana-Champaign; Stony Brook University; National Chengchi University(伊利诺伊大学厄巴纳-香槟分校; 纽约州立大学石溪分校; 国立政治大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过受控实验考察大型语言模型能否识别序列结构中的潜在策略与规则,发现长上下文无益于识别,且高阶依赖显著降低规则恢复,表面保真可能掩盖错误机制。

AI 中文摘要

大型语言模型(LLMs)越来越多地被用作交互式智能体和模拟器,但目前尚不清楚它们能否恢复表面动作频率之外的潜在序列结构。这一区别对于行为模拟至关重要,因为在行为模拟中,动作往往由先前情境而非边际频率单独塑造。我们通过受控的双人石头剪刀布互动和单人随机n-gram延续任务来研究这一问题。在这些实验中,我们测试LLMs能否识别潜在策略、遵循简单马尔可夫规则以及维持高阶条件依赖。我们的框架将分布匹配与条件规则遵循区分开来。结果表明,更长的上下文并不能提高识别能力,正确的识别并不能确保忠实的模拟,而高阶依赖会显著降低规则恢复能力。因此,表面上的行为保真度可能掩盖错误的生成机制。

英文摘要

Large language models (LLMs) are increasingly used as interactive agents and simulators, yet it remains unclear whether they can recover latent sequential structure beyond surface action frequencies. This distinction is critical for behavioral simulation, where actions are often shaped by prior context rather than marginal frequencies alone. We study this question using controlled two-player Rock--Paper--Scissors interactions and a one-player stochastic n-gram continuation task. Across these experiments, we test whether LLMs can identify latent strategies, follow simple Markov rules, and sustain higher-order conditional dependencies. Our framework separates distribution matching from conditional rule following. Results show that longer context does not improve identification, correct recognition does not ensure faithful simulation, and higher-order dependencies substantially degrade rule recovery. Apparent behavioral fidelity can therefore mask incorrect generative mechanisms.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑