代理何时应检查外部状态?为存储意图的观察进行预算分配
When Should Agents Check External State? Budgeting Observations for Stored Intentions
浏览论文内容
中文总结 AI 辅助
针对前瞻记忆代理中存储意图的外部状态检查,提出BudgetPM框架,在共享预算下分配观察资源,通过静态评分和序列蒸馏策略,显著减少观察次数并保持质量,得出需求-容量设计规则。
中文摘要 AI 辅助
前瞻记忆使代理能够保留与未来条件相关联的意图,但存储的意图并不揭示该条件当前是否成立。检查该条件可能需要访问网络、多步骤工具使用以及付费调用。现有系统决定意图何时需要关注,但并未在共享预算下分配由此产生的观察。我们首次提出了在共享情节预算下,存储意图所需外部观察的资源分配公式。BudgetPM提供了两种策略变体,它们共享一个硬预算执行器。BudgetPM-Static使用轻量级Logistic评分器来学习检查是否改善当前决策。BudgetPM-Sequential将完整情节的事后调度蒸馏为轻量级策略,该策略在部署时仅使用查询前信息来决定何时花费或保留容量。我们将BudgetPM与两个公共记忆代理系统、五个匹配对照以及四个手工设计的监控或预算适应规则进行了评估。在两个基准和三个骨干网络上,BudgetPM-Static优于改编的Mem0和PMA工作流。在PM-Bench上,其Logistic评分器达到了与更高容量评分器相当的质量-成本操作点,并以减少42-54%的观察保留了99.9-100%的无约束质量。在严重稀缺和相同硬上限下,BudgetPM-Sequential比最强测试的自然监控调度高出1.92-2.58个Set F1点。它以减少16-33%的观察达到相同的Set F1和按时召回率。匹配归因、精确成本分析和固定预算负载干预将此增益与当前和未来机会之间的竞争联系起来。这些结果产生了需求-容量设计规则:当容量覆盖需求时,局部门控有效,而当观察跨时间竞争时,未来感知监督增加价值。
英文摘要
Prospective memory allows an agent to retain an intention tied to a future condition, but the stored intention does not reveal whether that condition currently holds. Checking it may require web access, multi-step tool use, and paid calls. Existing systems decide when intentions require attention, but do not allocate the resulting observations under a shared budget. We introduce the first resource-allocation formulation for the external observations required by stored intentions under a shared episode budget. BudgetPM offers two policy variants that share a hard-budget executor. BudgetPM-Static uses a lightweight Logistic scorer to learn whether a check improves the current decision. BudgetPM-Sequential distills full-episode hindsight schedules into a lightweight policy that decides when to spend or reserve capacity using only pre-query information at deployment. We evaluate BudgetPM against two public memory-agent systems, five matched controls, and four hand-designed monitoring or budget-adaptation rules. Across two benchmarks and three backbones, BudgetPM-Static outperforms adapted Mem0 and PMA workflows. On PM-Bench, its Logistic scorer reaches competitive quality--cost operating points alongside higher-capacity scorers and retains 99.9--100\% of unconstrained quality with 42--54\% fewer observations. Under severe scarcity and the same hard caps, BudgetPM-Sequential exceeds the strongest tested natural monitoring schedule by 1.92--2.58 Set F1 points. It reaches the same Set F1 and on-time recall with 16--33\% fewer observations. Matched attribution, exact-cost analysis, and a fixed-budget load intervention link this gain to competition between present and future opportunities. These results yield a demand--capacity design rule: local gating works when capacity covers demand, while future-aware supervision adds value when observations compete across time.