AI 中文总结
本文提出执行信息需求(EIR)及LACUNA和VESTIGE框架,用于评估有状态智能体在长期任务中保留和恢复关键信息的能力,实验表明恢复缺失结果可显著提升准确率。
AI 中文摘要
长期运行的智能体必须保留后续步骤所依赖的信息。我们提出了执行信息需求(EIR),这是在特定任务和访问条件下,为正确完成任务而必须保持可访问的信息的下限。我们开发了LACUNA框架,该框架生成具有已知依赖关系的任务,并将信息需求、保留和恢复与单个操作的难度分开变化。在四个模型上,恢复缺失结果可将受影响回忆步骤的准确率提高到100%,而等长不相关信息的准确率为0%。仅靠足够的存储并不能确保成功:保留策略可能会丢弃所需结果,错误可能会通过后续计算传播,智能体可能会在恢复完成前停止。我们还引入了VESTIGE,它利用智能体执行轨迹构建语义图,并衡量真实任务的信息需求。在72,562条软件智能体轨迹中,VESTIGE揭示了失败运行中与解决方案相关的重读的距离相关下降更陡峭(每次距离加倍,相对风险为0.951),而调整后的峰值需求单独与失败无关。总之,这些贡献支持评估智能体是否保留和恢复其任务所需的信息。
英文摘要
Long-running agents must preserve information that later steps depend on. We introduce the Execution Information Requirement (EIR), a lower bound on the information that must remain accessible for correct completion under specified task and access conditions. We develop LACUNA, a framework that generates tasks with known dependencies and varies information demand, retention, and recovery separately from the difficulty of individual operations. Across four models, restoring a missing result raises accuracy on affected recall steps to 100%, compared with 0% for equal-length irrelevant information. Sufficient storage alone does not ensure success: retention policies can discard required results, errors can propagate through later computations, and agents can stop before recovery is complete. We also introduce VESTIGE, which uses agent execution traces to construct semantic graphs and measure information demand for real tasks. Across 72,562 software-agent trajectories, VESTIGE reveals a steeper distance-related decline in solution-relevant rereading for failed runs (RR 0.951 per distance doubling), while adjusted peak demand alone is not associated with failure. Together, these contributions support evaluating whether agents preserve and recover the information their tasks require.
Comments35 pages, 12 figures, 15 tables. Supplementary material included