arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.14138cs.LGcs.AI

LIMBO:面向LLM智能体的终身推理时记忆与预算优化

LIMBO: Lifelong Inference-Time Memory and Budget Optimization for LLM Agents

  • University of California, San Diego(加利福尼亚大学圣迭戈分校)
  • West Virginia University(西弗吉尼亚大学)

机构由 AI 辅助整理,请以论文原文为准。

Siddharth Sharma, Nilesh Prasad Pandey, Onat Gungor, Tajana Rosing

AI总结:

LIMBO提出首个在线框架,将记忆视为可控推理时资源,联合优化终身LLM智能体的记忆策略与推理预算,在LifelongAgentBench上以更低成本实现更优或相当的性能。

AI中文摘要:

随着LLM智能体被集成到日益复杂的工作流程中,它们必须持续获取新能力,同时保持对先前学习任务的胜任力。终身智能体通过经验回放来解决这一问题,即将过去的交互注入提示中,以便在推理时利用先前的经验。然而,回放并非没有代价:每条回放的轨迹都与检索、推理、工具使用和验证竞争同一有限的提示和计算预算,这使得有效的资源分配至关重要。现有方法使用固定的回放策略来分配这些资源,而不考虑回放对当前任务是否有益。我们将此识别为推理时记忆分配,这是终身智能体面临的一个独特问题类别,并引入LIMBO:据我们所知,这是第一个将记忆视为可控推理时资源并为每个传入任务联合优化记忆策略和推理预算的在线框架。与先前固定回放策略或需要模型权重、教师监督或离线重训练的方法不同,LIMBO以单次遍历的方式在线学习这种分配,明确平衡任务性能和推理成本,而无需修改底层智能体。在LifelongAgentBench上的三个LLM骨干网络上,LIMBO实现了比最先进的记忆增强基线更好的成本-准确性权衡,并且以高达约83%更低的推理成本(平均约53%)几乎匹配所有最强的此类基线。LIMBO无需重训练即可跨模型和环境调整其策略,表明有效的分配可以在线学习而非手动指定。

英文摘要:

As LLM agents become integrated into increasingly complex workflows, they must continually acquire new capabilities while retaining competence on previously learned tasks. Lifelong agents address this through experience replay, injecting past interactions into the prompt to leverage prior experience during inference. However, replay is not free: every replayed trajectory competes with retrieval, reasoning, tool use, and verification for the same limited prompt and compute budget, making effective resource allocation essential. Existing approaches allocate these resources using fixed replay policies, regardless of whether replay is beneficial for the current task. We identify this as inference-time memory allocation, a distinct problem class for lifelong agents, and introduce LIMBO: the first online framework to our knowledge that treats memory as a controllable inference-time resource and jointly optimizes memory strategy and inference budget for each incoming task. Unlike prior approaches that fix the replay policy or require model weights, teacher supervision, or offline retraining, LIMBO learns this allocation online in a single pass, explicitly balancing task performance and inference cost without modifying the underlying agent. Across three LLM backbones on LifelongAgentBench, LIMBO achieves better cost-accuracy tradeoffs than state-of-the-art memory-augmented baselines and nearly matches all strongest such baselines at up to ~83% lower inference cost (~53% on average). LIMBO adapts its policy across models and environments without retraining, demonstrating that effective allocation can be learned online rather than manually specified.

补充信息

↑