ARSM:用于智能体推理压缩的自回归状态机
ARSM: Auto-Regressive State Machine for Agentic Reasoning Compression
AI总结:
提出ARSM,一种无需训练的轻量级框架,通过结构化状态演化压缩智能体推理,在保持任务性能的同时减少令牌消耗。
AI中文摘要:
尽管基于大型语言模型(LLM)的智能体通过将推理与外部环境交互交错进行,在长周期任务中展现出强大的能力,但上下文的持续累积迅速造成了关键的内存瓶颈。现有的内存压缩方法依赖于特定任务的优化或外部辅助模型,引入了显著的计算开销。此外,由此产生的压缩表示往往会丢失结构化关系,导致信息稀释、注意力崩溃以及决策一致性下降。为解决这些局限,我们提出了自回归状态机(ARSM),一种轻量级、无需训练的框架,通过结构化状态演化实现原位推理压缩。ARSM引入了两个关键组件:(i)轨迹抽象机制,将交互历史重组为紧凑的假设-行动-结果(HAR)微链;(ii)动态状态机,通过原子操作和压缩控制参数来调节分层内存。这些组件统一在一个自回归、自压缩的生成空间中,其中每个模型输出同时执行外部行动执行和内部状态更新。我们在Webshop、多目标多跳问答和SWE-Bench Lite数据集上评估了ARSM。实验结果表明,ARSM在保持任务性能的同时减少了令牌消耗,为长周期任务中的可扩展自主智能体提供了一条实用且经济高效的路径。
英文摘要:
While Large Language Model (LLM)-based agents demonstrate strong capabilities in long-horizon tasks by interleaving reasoning with external environment interactions, the continuous accumulation of context rapidly creates a critical memory bottleneck. Existing memory compression methods rely on task-specific optimization or external auxiliary models, introducing significant computational overhead. Furthermore, the resulting compressed representations tend to lose structured relationships, leading to information dilution, attention collapse, and degraded decision consistency. To address these limitations, we propose Auto-Regressive State Machine (ARSM), a lightweight training-free framework that enables in-situ reasoning compression through structured state evolution. ARSM introduces two key components: (i) a trajectory abstraction mechanism that reorganizes interaction histories into compact Hypothesis-Action-Result (HAR) micro-chains; (ii) a dynamic state machine that regulates hierarchical memory through atomic operations and a compression-control parameter. These components are unified within an auto-regressive, self-compressive generation space, where each model output jointly performs external action execution and internal state updates. We evaluate ARSM on Webshop, Multi-Objective Multi-Hop QA, and SWE-Bench Lite datasets. Experimental results show that ARSM maintains the task performance while simultaneously reducing token consumption, offering a practical, cost-effective route toward scalable autonomous agents for long-horizon tasks.