arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.05988cs.LGcs.AI

DART:用于算法学习的分布对抗循环训练

DART: Distributional Adversarial Recurrent Training for Algorithm Learning

  • FPT IS AI R&D Center(FPT信息系统人工智能研发中心)
  • VNU – University of Engineering and Technology(越南国立大学工程技术大学)

机构由 AI 辅助整理,请以论文原文为准。

Hieu Tran Bao, Phung Thanh Dang, Pham Quang Nhat Minh, Hoang Thanh Tung

AI总结:

DART提出用局部目标分布替代单点监督,通过对抗训练提升循环推理模型在算法学习中的解质量、稳定性和鲁棒性,在迷宫、国际象棋和数独任务上验证有效。

AI中文摘要:

循环推理模型(RRMs)能够解决结构化问题,通过在隐藏空间中的迭代计算实现从易到难的泛化。这些模型通常使用实例级监督进行训练,但随着任务难度的增加,这种监督方式问题日益凸显:有效解仅占据解空间中的极小区域,而无效解则迅速激增。我们提出了分布对抗循环训练(DART),这是一种训练框架,它用围绕真实解周围的局部目标分布替代单点监督,并通过对抗目标使模型输出与该分布对齐。DART提供了更丰富的学习信号,并鼓励更稳定的迭代轨迹朝向有效解。在迷宫、国际象棋和带掩码的数独任务上,使用多种循环推理模型(包括深度思考系统和微型递归模型)进行评估时,DART在评估的分布偏移下提高了解的质量、稳定性和鲁棒性。与标签平滑、高斯软化目标和渐进式训练的比较表明,DART的效果不能仅用目标软化来解释,并且与稳定长时程循环的训练方案互补。这些结果将DART确定为一种在评估的循环推理模型中提高鲁棒性的有前景的方法。

英文摘要:

Recurrent reasoning models (RRMs) can solve structured problems, achieving easy-to-hard generalization through iterative computation in hidden space. These models are typically trained with instance-level supervision, which becomes increasingly problematic as task difficulty grows: valid solutions occupy a tiny region of the solution space, while invalid solutions proliferate rapidly. We propose Distributional Adversarial Recurrent Training (DART), a training framework that replaces single-point supervision with a local target distribution around the ground-truth solution and aligns model outputs with this distribution through an adversarial objective. DART provides a richer learning signal and encourages more stable iterative trajectories toward valid solutions. When evaluated on Maze, Chess, and masked Sudoku with multiple RRMs, including Deep Thinking Systems and Tiny Recursive Models, DART improves solution quality, stability, and robustness under the evaluated distribution shifts. Comparisons with label smoothing, Gaussian softened targets, and progressive training show that DART is not explained by target softening alone and is complementary to training schemes that stabilize long-horizon recurrence. These results identify DART as a promising approach for improving robustness across the evaluated recurrent reasoning models.

↑