arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.19523cs.CL

当推理缩小行动范围:大语言模型游戏中的多样性崩溃

When Reasoning Narrows the Move: Diversity Collapse in LLM Game Play

Junyi Sha, Renfei Tan, David Simchi-Levi

首次发表
浏览论文内容

中文总结 AI 辅助

研究大语言模型游戏中监督微调对行为多样性的影响,发现推理模式生成常抑制行动多样性,标准SFT会致过早多样性崩溃,行动增强可部分缓解,指出窄支持模仿是策略崩溃源,SFT中保持行动支持对维持探索行为很重要。

中文摘要 AI 辅助

监督微调(SFT)被广泛用于使大语言模型适应下游任务,但其对序列决策中行为多样性的影响仍未得到充分探索。我们在基于井字棋变体的确定性棋盘游戏的受控套件中研究了这个问题,其中最优行动是可精确计算的,并且多样性可以直接测量。通过状态级评估、竞技场游戏玩法和训练轨迹,我们发现推理模式生成经常会抑制行动多样性,而不会统一提高行动准确性。此外,标准的SFT提高了准确性,但经常会导致过早的多样性崩溃,这超过了准确性-多样性权衡所要求的最低限度。然后我们表明,行动增强,即在每个状态的所有最优行动上进行训练而不是单个示范行动,将部分减轻这种影响。我们的结果将窄支持模仿确定为大语言模型决策中策略崩溃的一个来源,并表明在SFT期间保持行动支持对于维持探索行为很重要。

英文摘要

Supervised fine-tuning (SFT) is widely used to adapt large language models to downstream tasks, but its effect on behavioral diversity in sequential decision-making remains under-explored. We study this question in a controlled suite of deterministic board games based on tic-tac-toe variants, where optimal actions are exactly computable and diversity can be measured directly. Across state-level evaluation, arena gameplay, and training trajectories, we find that reasoning-mode generation frequently suppresses action diversity without uniformly improving action accuracy. Furthermore, standard SFT improves accuracy but often induces premature diversity collapse, which exceeds what is minimally required by the accuracy-diversity tradeoff. We then show that action augmentation, which trains on all optimal actions per state rather than a single demonstrated action, would partially mitigates this effect. Our results identify narrow-support imitation as a source of policy collapse in LLM decision-making and suggest that preserving action support during SFT is important for maintaining exploratory behavior.

发表机构

  • Institute for Data, Systems, and Society(数据、系统与社会研究所)

机构由 AI 辅助整理,请以论文原文为准。

↑