AI 中文总结
SCLATE是一个统一事件调度与回滚引擎的基础设施,支持持续学习智能体的训练评估;实验表明附加记忆系统并非总优于原生记忆,且后训练Qwen3.5-4B显著提升性能。
AI 中文摘要
持续学习智能体是由模型、框架(harness)和记忆(memory)组成的系统,在长时间的多会话(multi-session)时间跨度上运行。对它们进行评估和训练需要将任务与智能体侧事件(如会话停止和开始、定时任务(crons)以及记忆整合)交错进行。然而,现有的基准测试和训练框架仅调度基准自身的事件,使得每个基准与智能体的组合都需要构建自定义的调度循环。我们提出了SCLATE,一个执行基础设施(substrate),在其中基准测试和未经修改的智能体通过适配器(adapter)将各自的事件添加到一个开放的事件调度器中。一个混合模拟时钟在共享时间线上运行这些事件,在智能体工作时实时流动,并跳过空闲间隙,从而将一个月长的场景压缩为数小时。SCLATE还充当一个回滚(rollout)引擎,可运行任何智能体的框架和记忆而无需修改,并通过容器内代理记录每次模型调用的令牌(tokens)和对数概率(log probabilities)。我们将七个基准移植到SCLATE,并在十个模型上对十种未经修改的框架和记忆配置进行了头对头比较。比较结果表明,添加的记忆系统并不总是可靠地优于框架的原生记忆,并且不同模型在使用相同框架和记忆时差异很大。随后,我们通过未经修改的框架和记忆系统对Qwen3.5-4B进行了后训练(post-train)。该模型学会了同时使用两者,读取的文件行数减少了6.8倍,SWE-bench Verified通过率提高了16.7个百分点,并写出了更丰富的记忆记录,同时其保留的MetaClaw准确率最高提升了11.8个百分点。
英文摘要
Continual-learning agents are systems of models, harnesses, and memory operating over long multi-session horizons. Evaluating and training them requires interleaving tasks with agent-side events such as session stop and start, crons, and memory consolidation. Yet existing benchmarks and training frameworks schedule only the benchmark's own events, leaving each benchmark and agent pair to build a custom scheduling loop. We present SCLATE, an execution substrate where benchmarks and unmodified agents each add their events to one open event scheduler through an adapter. A hybrid simulated clock runs these events on a shared timeline, flowing in real time while the agent works and skipping idle gaps, which compresses a month-long scenario into hours. SCLATE also serves as a rollout engine that runs any agent's harness and memory unmodified, recording the tokens and log probabilities of every model call through an in-container proxy. We port seven benchmarks to SCLATE and compare ten unmodified harness and memory configurations head to head on ten models. The comparison shows that an added memory system does not reliably beat the harness's native memory and that models differ widely in how they use the same harness and memory. We then post-train Qwen3.5-4B through unmodified harnesses and memory systems. The model learns to use both, reading 6.8x fewer file lines with a 16.7-point higher SWE-bench Verified pass rate, and writing richer memory records, while its held-out MetaClaw accuracy rises by up to 11.8 points.