发表机构
Alibaba Group; Columbia University(阿里巴巴集团; 哥伦比亚大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出名为Scroll的上下文管理器,将智能体会话视为可执行环境,结合事件日志与Python内核实现程序化上下文管理,在多个长程智能体基准测试中大幅超越现有最优系统。
AI 中文摘要
大语言模型(LLM)智能体越来越多地承担长时运行任务,其历史交互内容远超单个模型上下文窗口的容量。现有方法要么压缩早期交互,要么将选定信息提取为固定的记忆表示,在未来需求明确前就确定了要保留的内容。我们提出Scroll,一种将每个智能体会话视为可执行会话环境的上下文管理器。该环境由仅追加的事件日志和沙箱式持久化Python内核支撑。内核在模型调用间维护类型化命名空间,允许将工具输出、检索到的历史和派生状态绑定到变量,而非在每次调用时序列化到提示词中。模型编写的代码通过exec搜索、具体化和转换会话状态;只有显式打印的投影会进入下一次调用的模型工作视图。因此,上下文管理成为一项编程任务,继承了LLM不断提升的编码能力,同时事件日志保留了无损的历史真实值。当工作视图接近其预算时,陈旧片段会被逐出但仍可恢复:逐出索引会保留与事件日志确切地址绑定的紧凑地标,使智能体可直接导航到已逐出的区域,而非搜索整个日志。以Qwen3.8-Max为骨干,Scroll在LongMemEval_S上达到94.8%;在BEAM_10M上达到73.1%,比已发表的最佳记忆系统高出5.1个百分点;在LOCA_256K上达到86.7%,比已发表的最佳长程智能体高出37.4个百分点。
英文摘要
LLM agents increasingly take on long-running tasks whose history grows far beyond a single model context window. Existing approaches compress earlier interactions or extract selected information into fixed memory representations, committing to what to preserve before future needs are known. We present Scroll, a context manager that treats each agent session as an executable Session Environment. The environment is backed by an append-only Event Log and a sandboxed, persistent Python kernel. The kernel maintains a typed namespace across model calls, allowing tool outputs, retrieved history, and derived state to be bound to variables rather than serialized into the prompt at each call. Model-written code searches, materializes, and transforms session state through exec; only explicitly printed projections enter the model's working view for the next call. Context management thus becomes a programming task that inherits the improving coding abilities of LLMs, while the Event Log preserves lossless historical ground truth. As the working view approaches its budget, stale spans are evicted but remain recoverable: an eviction index keeps compact landmarks tied to exact Event Log addresses, so that the agent navigates directly to evicted regions instead of searching the full log. With Qwen3.8-Max as the backbone, Scroll achieves 94.8% on LongMemEval_S; 73.1% on BEAM_10M, surpassing the best published memory system by 5.1 points; and 86.7% on LOCA_256K, exceeding the best published long-horizon agent by 37.4 points.