arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

StateTree:通过强化学习增强长期对话推理

StateTree: Enhancing Long-Term Dialogue Reasoning via Reinforcement Learning

Naen Xu, Wanqing Cui, Yibo Hu, Shixin Hong, Hengyu An, Meiguang Jin, Junfeng Ma, Tianyu Du

arXiv 2609.38809首次发表:更新:

发表机构

Zhejiang University; Taobao & Tmall Group of Alibaba(浙江大学; 阿里巴巴淘宝天猫集团)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出StateTree强化学习方法,通过树状路径追踪任务增强长期对话推理,在10K上下文训练后泛化至128K,显著提升跨会话检索与多跳推理性能。

AI 中文摘要

作为个性化助手部署的大型语言模型必须对长期且不断演变的交互历史进行推理。然而,在长期对话推理中,相关证据分散在各个会话中,偏好可能随时间被修订,并且标准的长上下文训练在数据稀缺和计算成本高昂的情况下无法解决这些挑战。我们提出StateTree,一种数据驱动的强化学习方法,它从具有可验证真实答案的稀缺对话中构建一个具有挑战性的辅助任务。StateTree通过一种树状路径追踪任务增强多会话对话:键值记录被嵌入到各个会话中,形成一棵二叉树。解决该任务要求模型通过跨会话检索记录并比较时间戳来解决分支,从而从根节点遍历到叶节点,然后在干扰叶节点中恢复隐藏的目标问题。我们应用课程强化学习训练,逐步增加树的深度,并引入一种组合变体,其边携带逐步推理片段,训练模型将部分线索组合成连贯的查询。在10K令牌上下文的训练下,StateTree泛化到128K令牌,无需全长度强化学习的成本,并展现出包括跨会话检索、时间推理、知识更新和组合多跳推理在内的能力。StateTree在保持短上下文一般推理能力的同时,优于基于监督微调和强化学习的基线。StateTree-7B在LongMemEval(128k)上取得了高达+23.60%的提升,StateTree-14B在LongMemEval上达到59.00%的准确率,超过了QwenLong-L1-32B(45.20%)。

英文摘要

Large language models deployed as personalized assistants must reason over long, evolving interaction histories. However, in long-term dialogue reasoning, relevant evidence is scattered across sessions, preferences may be revised over time, and standard long-context training fails to address these challenges under data scarcity and prohibitive computational costs. We propose StateTree, a data-driven RL method that constructs a challenging auxiliary task from scarce dialogues with verifiable ground truth. StateTree augments multi-session dialogues with a tree-structured path-tracing task: key-value records are embedded across sessions to form a binary tree. Solving the task requires the model to traverse from root to leaf by retrieving records across sessions and comparing timestamps to resolve branches, then recover the hidden target question among distractor leaves. We apply curriculum RL training progressively increasing tree depth and introduce a compositional variant whose edges carry step-level reasoning fragments, training the model to compose partial cues into coherent queries. Trained on 10K-token contexts, StateTree generalizes to 128K tokens without full-length RL costs and exhibits capabilities including cross-session retrieval, temporal reasoning, knowledge update, and compositional multi-hop reasoning. StateTree outperforms both SFT and RL-based baselines while preserving short-context general reasoning. StateTree-7B achieves gains up to +23.60% on LongMemEval (128k), and StateTree-14B reaches 59.00% accuracy on LongMemEval, surpassing QwenLong-L1-32B (45.20%).

CommentsNeurIPS 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑