arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于动作中心图的推理时缩放与情景记忆的衔接

Bridging Inference-Time Scaling and Episodic Memory with Action-Centric Graphs

Xu Zheng, Chaohao Lin, Zhuomin Chen, Weijieying Ren, Haifeng Chen, Wei Cheng, Dongsheng Luo

arXiv 2607.27415首次发表:更新:

发表机构

Florida International University; Stanford University; NEC Laboratories America; Singapore Management University(佛罗里达国际大学; 斯坦福大学; 美国NEC实验室; 新加坡管理大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出GAMER框架,通过动作中心图衔接推理缩放与情景记忆,采用双流时间差分学习优化决策,在多基准上较普通基线提升了20.81%的成功率与6.17%的进度率。

AI 中文摘要

推理时缩放的最新进展已极大释放了大语言模型(LLMs)的复杂推理能力,但对于智能体而言,这些方法存在严重低效问题——以无状态方式运行并产生冗余搜索过程。现有记忆机制大多依赖LLMs的推理能力,导致计算成本过高。本文提出一种新型框架GAMER(基于动作中心图的情景推理记忆,Graph-based Action-centric Memory with Episodic Reasoning),用于衔接推理缩放与情景记忆的缺口。该方法将历史推理建模为动态动作中心图,通过将记忆机制与LLMs解耦,相比记忆机制基线提供更少的记忆上下文,从而节省token使用量与资金。为有效从图中提取知识,采用双流时间差分学习机制,基于过往成功与失败估计动作节点的正向(建议)和负向(规避)价值。在推理阶段,该学习到的价值函数双向优化决策,正向价值提供动作建议,负向价值指示高风险动作;通过在图上执行高效搜索,该方法显著提升推理缩放的效率。在多个基准上的实验表明,GAMER相比普通基线在成功率和进度率上分别实现了20.81%和6.17%的优越性能。

英文摘要

Recent advancements in inference-time scaling have significantly unlocked the complex reasoning capabilities of Large Language Models~(LLMs). However, for agents, these approaches suffer from a critical inefficiency, operating in a stateless manner and engaging in redundant search processes. Existing memory mechanisms largely rely on the reasoning capabilities of LLMs, leading to prohibitive computational costs. In this paper, we propose a novel framework, \textit{GAMER}~(Graph-based Action-centric Memory with Episodic Reasoning), that bridges the gap between inference scaling and episodic memory. Our approach models historical reasoning as a dynamic \textit{Action-Centric Graph}. By decoupling the memory mechanism from LLMs, our method can save token/money usage by providing less memory context than memory mechanism baselines. To extract knowledge from the graph effectively, we use a dual-stream Temporal Difference learning mechanism to estimate the positive~(suggestion) and negative~(avoidance) value of action nodes based on past successes and failures. During the inference phase, this learned value function optimizes decision-making bi-directionally, so that positive values provide action suggestions, while negative values indicate high-risk actions. By performing efficient searches on the graph, our method significantly improves the efficiency of inference scaling. Experiments on multiple benchmarks demonstrate that \textit{GAMER} achieves superior performance by \textbf{20.81\%/6.17\%} for success/progress rate compared to vanilla baselines.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑