发表机构
Ant International, Ant Group(蚂蚁国际,蚂蚁集团)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
FlowState将执行状态作为可重访的记忆,通过增量状态更新和渐进式状态访问,在长时程任务中提升性能并降低token消耗。
AI 中文摘要
长时程任务要求LLM智能体持续利用早期交互中的信息。然而,保留完整历史会增加上下文成本,而压缩历史则可能丢失后续所需细节,且历史信息的相关性往往随着任务进展才变得明显。为应对这些挑战,我们提出FlowState,它将执行状态视为可在请求间保留和重访的记忆,将当前决策与历史信息的重用统一起来。FlowState保留语义类型化的状态节点、它们之间的关系以及对原始工具观测的引用,将持久保留与按需访问分离。在单个执行循环内,增量状态更新(ISU)基于新输入和反馈维护当前状态,而渐进式状态访问(PSA)在推理过程中按需逐步揭示历史状态和支持证据。这些机制共同使智能体能够根据新信息重新评估先前决策并指导后续行动。与使用相同DeepSeek-V4-Flash模型的完整上下文基线相比,FlowState在MemoryArena上的平均成功率和τ³-Bench上的平均通过率分别提高了4.55和13.95个百分点,同时总token消耗分别减少了43.2%和40.6%。这些结果证明了FlowState在长时程任务上的性能和效率优势。
英文摘要
Long-horizon tasks require LLM agents to continually draw on information from earlier interactions. However, retaining the full history increases context costs, while compressing it risks losing details needed later, and the relevance of historical information often becomes apparent as the task progresses. To address these challenges, we propose FlowState, which treats execution state as memory that can be retained and revisited across requests, unifying current decision-making with the reuse of historical information. FlowState preserves semantically typed state nodes, their relations, and references to raw tool observations, separating persistent retention from on-demand access. Within a single execution loop, Incremental State Update (ISU) maintains the current state based on new inputs and feedback, while Progressive State Access (PSA) progressively reveals historical states and supporting evidence as needed during reasoning. Together, these mechanisms enable agents to reassess prior decisions in light of new information and guide subsequent actions. Compared with a full-context baseline using the same DeepSeek-V4-Flash model, FlowState improves the average success rate on MemoryArena and the average pass rate on $τ^3$-Bench by 4.55 and 13.95 percentage points, respectively, while reducing total token consumption by 43.2% and 40.6%. These results demonstrate the performance and efficiency advantages of FlowState on long-horizon tasks.