AI 中文总结
研究大语言模型智能体运行后磁盘持久数据评估问题,引入AgentFootprint基准测试,通过序列化感知度量套件测量多方面指标,解决测量陷阱,给出不同配置差异及存储优化结果,确立持久存储为资源指标。
AI 中文摘要
大语言模型智能体基准测试衡量任务完成情况、可靠性和推理成本,但不衡量智能体运行在磁盘上留下的持久数据,包括日志、上下文快照、检查点和调试跟踪。我们引入了AgentFootprint,这是一个运行后智能体存储足迹的跨框架基准测试。其序列化感知度量套件测量总保留量、通道组成、重复率、增长、可压缩性和对话历史可重构性。它解决了一个测量陷阱:由于数据库分页和JSON转义会掩盖重复内容,简单的字节级测量会使重复率低估一个数量级。固定跟踪控制将智能体生成的逻辑卷与持久层放大分开:通过七个持久化框架重放相同轨迹会产生6.7倍的差异。在相同模型、工具和任务下,准确率为100%的配置在保留字节数上相差15.7倍,尽管它们的默认设置支持不同的恢复和审计功能。三种全历史配置在重复观察压力任务上呈超线性增长。来自108个实例归一化的SWE-bench验证提交的导出轨迹每个实例跨越三个数量级,与解析率没有可检测到的相关性。内容寻址存储在保留每个可重构性分数的同时,将保留量减少了4.8倍至32.7倍。这些结果将持久存储确立为与准确率和可重构性一起报告的资源指标。
英文摘要
LLM agent benchmarks measure task completion, reliability, and inference cost, but not the persistent data an agent run leaves on disk, including logs, context snapshots, checkpoints, and debug traces. We introduce AgentFootprint, a cross-framework benchmark of post-run agent storage footprint. Its serialization-aware metric suite measures total retention, channel composition, duplication, growth, compressibility, and conversation-history reconstructability. It addresses a measurement trap: naive byte-level measurement understates duplication by an order of magnitude because database paging and JSON escaping obscure repeated content. A fixed-trace control separates agent-generated logical volume from persistence-layer amplification: replaying the same trajectory through seven persisting frameworks yields a 6.7x spread. Under identical models, tools, and tasks, configurations with 100% accuracy differ by 15.7x in retained bytes, although their defaults support different recovery and audit capabilities. Three full-history configurations grow superlinearly on a repeated-observation stress task. Exported trajectories from 108 instance-normalized SWE-bench Verified submissions span three orders of magnitude per instance, with no detectable correlation with resolve rate. A content-addressed store reduces retention by 4.8x-32.7x while preserving every reconstructability score. These results establish persistent storage as a resource metric to report jointly with accuracy and reconstructability.
Comments17 pages, 5 figures; includes supplementary material