arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

活动帧:面向智能体记忆与回放的确定性屏幕活动编译

Activity Frames: Deterministic Screen-Activity Compilation for Agent Memory and Replay

Nossa Iyamu

arXiv 2608.05784首次发表:更新:

AI 中文总结

该研究提出确定性零模型流水线将屏幕活动编译为智能体记忆,可缩小上下文体积、提升问答准确率,还能获取智能体成本模型所需的例程开销比率与重复率等参数,相关资源开源。

AI 中文摘要

计算机使用智能体在重新推导用户已执行的例程时会付出全部前沿推理成本,因为当前智能体的记忆记录的是用户所说的内容,而非用户所做的内容。我们采用确定性的零模型流水线将被动捕获的屏幕活动编译为智能体记忆:该流水线将本地捕获流分割为类型化活动帧,这些帧是带有应用程序、站点、时间、输入量以及指向原始行的证据指针的有界片段,且流水线中无模型参与,因此输出是字节级相同、可缓存且可机械审计的。在一名专业人员的包含51个活跃日的单用户语料库(共128756帧)上,该编译器将一天的原始捕获数据缩减为提示就绪的上下文块,体积缩小86倍,耗时68毫秒;读取该块的智能体对当天问题的回答准确率达98.4%(Wilson 95%置信区间为91.7-99.7%),相较于同一捕获内容的LLM摘要的66-80%,中端模型读取该块的表现与前沿模型相当。该编译器还可作为需求侧成本工具:读取被动的、委托前的人类活动而非智能体部署,它提供了智能体成本模型所假设但据我们所知尚未测量的两个参数:例程开销比率R和例程重复率h。我们报告了R的首个值(建模上限为60-343倍),以及实际全舰队令牌上限约8%时的样本内可委托重复率9.0%、样本外7.7%;编译后的例程在模型外确定性回放,在防御匹配命中时零模型令牌的情况下得到现场演示。模式、编译器和评估工具均为开源。

英文摘要

Computer-use agents pay full frontier inference to re-derive routines their user has already performed, because an agent's memory today records what the user said, not what the user did. We compile passively captured screen activity into agent memory with a deterministic, zero-model pipeline: it segments a local capture stream into typed activity frames, bounded episodes carrying application, site, timing, input volume, and evidence pointers back to the raw rows, with no model in the loop, so the output is byte-identical, cacheable, and mechanically auditable. On one professional's single-user corpus of 128,756 frames over 51 active days, the compiler reduces a day of raw capture to a prompt-ready context block 86x smaller in 68 ms, and an agent reading that block answers questions about the day at 98.4% accuracy (Wilson 95% CI 91.7-99.7%) against an independent oracle, versus 66-80% for an LLM summary of the same capture, a mid-tier model reading the block matching a frontier one. The same compiler doubles as a demand-side cost instrument. Read off passive, pre-delegation human activity rather than agent rollouts, it supplies two parameters that agent-cost models assume but, to our knowledge, have not measured: the Routine Overhead Ratio R and the routine recurrence h. We report first values of R, a modeled upper bound, at 60-343x, and a delegable recurrence of 9.0% in-sample and 7.7% out-of-sample, for a realistic all-fleet token ceiling near 8%; a compiled routine replays deterministically with the model out of the loop, demonstrated live at zero model tokens on a guard-matched hit. Schema, compiler, and evaluation harness are open.

Comments14 pages, 5 figures, 4 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑