超越检索:基于查询条件的长程智能体轨迹复用
Beyond Retrieval: Query-Conditioned Reuse of Long-Horizon Agent Trajectories
浏览论文内容
中文总结 AI 辅助
该研究针对长程智能体轨迹检索后复用的瓶颈,构建评估框架并提出QCR方法,在多任务基准上验证其优于直接轨迹复用,实现检索质量与经验转化的分离。
中文摘要 AI 辅助
检索能够识别出可能相关的过往轨迹,但无法明确在用户、实体、约束或环境状态发生变化后,行动智能体应如何使用该轨迹。我们将这种检索后的复用步骤确定为长程轨迹记忆的一个独特瓶颈,并构建了一个评估框架,该框架在保持候选检索结果、目标状态、模型、解码方式和工具预算固定的同时,改变为智能体提供的支持。我们用基于查询条件的复用(Query-Conditioned Reuse, QCR)实例化该框架,QCR是一种刻意设计的简单目标绑定注释,记录可复用的流程、用于恢复的绑定、适用条件以及验证要求。QCR用于测试复用假设,而非声称其为普遍优选的记忆格式。在WebArena、WorkArena和AppWorld的2391个目标实例中,QCR达到了62.3%的平均成功率,比完整轨迹高出10.7个百分点,同时使用的在线代币减少了48.9%。摘要重排序为94.8%的目标选择了可复用记忆,使最终任务成功率与最优可复用选择器的差距在1.8个百分点以内。按轨迹长度和源-目标绑定偏移的分析显示,随着轨迹变长或源特定值发生变化,直接轨迹注入会丧失大部分效用,而目标绑定支持则能保留更大比例的观测增益。该框架将检索质量与将检索到的经验转化为新任务的安全、有用支持的问题分离开来。
英文摘要
Retrieval can identify a past trajectory that may matter, yet it does not specify how an acting agent should use that trajectory after users, entities, constraints, or environment state have changed. We identify this post-retrieval reuse step as a distinct bottleneck for long-horizon trajectory memory and formulate an evaluation framework that holds candidate retrieval, target state, model, decoding, and tool budget fixed while varying the support delivered to the agent. We instantiate the framework with query-conditioned reuse (QCR), a deliberately simple target-bound note that records a reusable procedure, bindings to recover, applicability conditions, and verification requirements. QCR serves to test the reuse hypothesis rather than to claim a universally preferred memory format. Across 2,391 target instances in WebArena, WorkArena, and AppWorld, QCR reaches 62.3% average Success, 10.7 points above Full Trajectory, while using 48.9% fewer online tokens. Summary reranking selects a reusable memory for 94.8% of targets, placing end-task Success within 1.8 points of an oracle reusable selector. Analyses by trajectory length and source--target binding shift show that direct trajectory injection loses much of its utility as traces grow longer or source-specific values change, whereas target-bound support preserves a larger share of the measured gain. The resulting framework separates retrieval quality from the problem of turning retrieved experience into safe, useful support for a new task.