发表机构
Max Planck Institute for Intelligent Systems; ELLIS Institute Tübingen; ETH Zurich; University of Oxford; Tübingen AI Center; Liquid AI(马克斯·普朗克智能系统研究所; ELLIS 蒂宾根研究所; 苏黎世联邦理工学院; 牛津大学; 蒂宾根人工智能中心; Liquid AI)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文证明两层Transformer参数化的流可解决图可达性,提出通过潜在展开训练和球面收缩稳定动态,使推理性能随积分步骤增加而提升,并在ProsQA、数独和迷宫任务上取得显著改进。
AI 中文摘要
流匹配能够在少量步骤内实现语言生成,但额外的积分步骤是否能提升推理能力仍不清楚。我们证明,由两层Transformer参数化的流可以解决图可达性问题,所需的积分步骤数量随目标距根节点的距离增加而增加。然而,标准的流语言模型在推理任务上可能无法从额外步骤中获益。我们将此限制归因于目标函数独立监督每个时间点,而未明确训练连续步骤以相互构建。为解决此问题,我们改为通过模型自身的潜在展开进行训练,在[0,1]的随机采样子区间上,仅在端点进行解码。在ProsQA上,这使准确率提升至97%,并使得性能随额外积分步骤的增加而提高。对于推理任务(如数独和迷宫)所需的更长展开,将潜在状态收缩到球面上可稳定动态,并在参数数量超过基线三倍以上的情况下取得显著收益。当与无参数的选择分数配对时,采样多个展开可进一步提升性能,尽管对于较长的答案,可靠的选择仍具挑战性。总之,这些结果为流推理建立了理论基础,并展示了展开训练、稳定的潜在动态和展开选择如何在实践中实现这种能力。
英文摘要
Flow matching enables language generation in few steps, but whether additional integration steps improve reasoning remains unclear. We prove that a flow parameterized by a two-layer Transformer can solve graph reachability, with the required number of integration steps increasing with the target's distance from the root. Yet, standard flow language models can fail to benefit from additional steps on reasoning tasks. We attribute this limitation to objectives that supervise each time point independently, without explicitly training successive steps to build on one another. To address this, we instead train through the model's own latent rollout over a randomly sampled subinterval of [0, 1], decoding only at the endpoint. On ProsQA, this raises accuracy to 97% and enables performance to improve with additional integration steps. For the longer rollouts required by reasoning tasks such as Sudoku and Maze, retracting the latent state onto a sphere stabilizes the dynamics and yields substantial gains over baselines with more than three times as many parameters. Sampling multiple rollouts further improves performance when paired with a parameter-free selection score, although reliable selection remains challenging for longer answers. Together, these results establish a theoretical basis for reasoning with flows and show how rollout training, stable latent dynamics, and rollout selection help realize this capacity in practice.
Comments23 pages, 5 figures, 7 tables