arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.09590cs.AIcs.RO

学习情境条件化思考策略以支持长期运行的LLM智能体

Learning Situation-Conditioned Thinking Policies for Long-Term LLM Agents

  • Chengdu University of Information Technology(成都信息工程大学)

机构由 AI 辅助整理,请以论文原文为准。

Hong Su

AI总结:

本文提出情境条件化思考记忆框架,将历史推理经验转化为轻量级策略以预测当前情境下的思考内容,实验证明其显著提升推理性能并大幅降低处理时间。

AI中文摘要:

长期运行的自主智能体必须重用累积的推理经验,同时避免显式历史记忆和LLM上下文无限增长。然而,现有的记忆机制主要检索、总结或压缩过去的内容,并不直接学习特定类型的思考应在何时被激活,也不从时间上分散的经验中发现新的思考知识。本文提出了一种情境条件化思考记忆框架,将历史推理经验转化为轻量级策略,用于预测在当前情境下应该思考什么,而将详细推理留给大型语言模型。情境可以表示时间或时空演化,而不仅仅是当前状态。临时经验还会跨多个独立情节定期分析,以识别重复的长期规律,这些规律被整合为新的思考知识,并进一步被轻量级策略内化。实验表明,学习到的策略在时间规则泛化上达到1.000的F1分数,将DeepSeek推理的F1分数从0.789提升至0.868,在30,000个历史情境下将每次查询的在线处理时间从0.3636毫秒降至0.0382毫秒,并在足够的重复跨经验证据后达到1.000的关系发现F1分数和未来思考准确性。

英文摘要:

Long-running autonomous agents must reuse accumulated reasoning experience without allowing explicit historical memory and LLM context to grow indefinitely. However, existing memory mechanisms mainly retrieve, summarize, or compress past content and do not directly learn when particular kinds of thinking should be activated or discover new thinking knowledge from temporally dispersed experiences. This paper proposes a situation-conditioned thinking memory framework that transforms historical reasoning experience into a lightweight policy for predicting what should be thought about in the current situation, while leaving detailed reasoning to a large language model. Situations may represent temporal or spatiotemporal evolution rather than only current states. Temporary experiences are also periodically analyzed across multiple independent episodes to identify repeated long-range regularities, which are consolidated into new thinking knowledge and further internalized by the lightweight policy. Experiments show that the learned policy achieves 1.000 F1 on temporal-rule generalization, improves DeepSeek reasoning F1 from 0.789 to 0.868, reduces online processing time from 0.3636 ms to 0.0382 ms per query at 30,000 historical situations, and reaches 1.000 relation-discovery F1 and future-thinking accuracy after sufficient repeated cross-experience evidence.

↑