询问工具,而非猜测:智能体工具调用掌握自身进度,服务系统应读取之
Ask the Tool, Don't Guess: Agent Tool Calls Hold Their Progress, and the Serving System Should Read It
- Tsinghua University(清华大学)
- Alibaba Cloud Computing(阿里云计算)
- Zhejiang University(浙江大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对智能体工具调用等待期间KV缓存管理依赖猜测的问题,提出让工具显式报告进度,经语料库验证信号可靠,接入生产引擎后显著降低TTFT。
AI中文摘要:
一个智能体请求在等待工具执行时耗费大量墙钟时间,其KV缓存在此期间持续占用GPU内存。服务系统通过猜测工具将运行多久来决定该缓存是保留、移除还是召回,猜测依据包括工具名称、历史记录、调用前声明的持续时间或引擎自身的占用率。我们证明,在调用开始前固定的任何估计都无法预知其实际持续时间,且此类估计甚至可能无法对调用进行排序。与此同时,正在运行的工具已持有答案,但智能体栈与工具共同将其静默。我们提议工具在运行期间明确报告其进度,并测量实现此目标所需的代价。对四个公开智能体语料库的普查发现,一旦进度被揭示,大多数工具时间中存在可读信号,其强度分为两种:剩余工作量的分数,或接近结束的准确信号。一个工具框架在不改变智能体所见内容的情况下恢复该信号,且对智能体的基准分数无可测量成本。在做出KV缓存决策的关键点上,所报告的进度比最佳已发布预测器准确数倍至一个数量级,且在环境变化时保持准确。通过少量小提示接入生产引擎,它将工具调用后首令牌时间(TTFT)的p90值相对于LRU降低了20.7%(仅HBM)和20.8%(HBM+DRAM),接近理想情况。服务系统不应猜测其工具能告知的信息。
英文摘要:
An agentic request spends substantial wall-clock time waiting for tools, and its KV cache holds GPU memory the whole time. Serving systems decide whether that cache stays, leaves, or comes back by guessing how long the tool will run, from the tool's name, its history, a duration declared before the call, or the engine's own occupancy. We show that no estimate fixed before a call starts can know its duration, and such estimates may not even rank the calls. Meanwhile, the running tool already holds the answer, but the agent stack together with the tool silences it. We propose that tool calls report their progress explicitly while they run, and we measure what that takes. A census of four public agent corpora finds a readable signal in most tool time once it is revealed, in two strengths: a fraction of the work remaining, or an accurate signal that the end is near. A harness recovers it without changing what the agent sees, at no measurable cost to the agent's benchmark score. At the points where a KV cache decision is made, the reported progress is between several times and an order of magnitude more accurate than the best published predictors, and it stays accurate when the environment changes. Plugged into a production engine through a few small hints, it cuts the p90 time to first token (TTFT) after a tool call by 20.7% (HBM only) and 20.8% (HBM + DRAM) against LRU, close to an oracle. A serving system should not guess what its tools can tell it.