arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.32283cs.DC

AgentReplay:令牌级轨迹重放对于公平的服务系统性能基准测试至关重要

AgentReplay: Token-Wise Trace Replay Is Essential for Fair Serving System Performance Benchmarking

  • University of California San Diego(加利福尼亚大学圣地亚哥分校)

机构由 AI 辅助整理,请以论文原文为准。

Zaifeng Pan, Michael Wang, Chris Wu, Zhengding Hu, Xinwei Qiang, Zhongkai Yu, Yufei Ding

AI总结:

针对LLM智能体服务中相同任务产生不同执行轨迹导致性能评估不公平的问题,提出AgentReplay框架,通过令牌级轨迹记录与重放消除工作负载变化,实现更公平的服务系统性能基准测试。

AI中文摘要:

基于大语言模型(LLM)的智能体执行多轮工作流,其中模型推理和工具调用交错进行,这使得高效的服务变得越来越重要。然而,评估服务优化具有挑战性,因为相同的任务可能产生不同的执行轨迹。生成的令牌的变化可能改变后续的提示、工具调用和推理轮次,这使得区分系统改进和工作负载变化变得困难。贪婪解码不能保证产生相同的输出,而仅重放序列长度会丢失影响前缀缓存和混合专家(MoE)路由的令牌信息。为了解决这些问题,我们提出了AgentReplay,一个可配置的用于智能体服务的轨迹记录和重放框架。AgentReplay记录输入/输出令牌、专家选择、请求依赖关系和工具持续时间。在重放期间,它强制使用记录的输出令牌,同时执行正常的自回归计算,并可选地控制MoE专家选择和工具延迟。这使得不同的系统能够执行相同的记录工作负载,同时保留它们自己的批处理、调度和并行化决策。我们进一步将轨迹生成与性能评估分离,使得兼容的较小模型能够重放使用更强大模型收集的长视野轨迹。我们的实验表明,AgentReplay中的令牌级重放有效地消除了贪婪解码和长度级重放无法避免的工作负载变化,从而在服务配置之间实现更公平的性能比较。

英文摘要:

LLM-based agents execute multi-turn workflows with interleaved model inference and tool calls, making efficient serving increasingly important. However, evaluating serving optimizations is challenging because identical tasks can produce different execution trajectories. Changes in generated tokens can alter subsequent prompts, tool calls, and reasoning turns, making it difficult to distinguish system improvements from workload variation. Greedy decoding does not guarantee identical outputs, while replaying only sequence lengths loses token information that affects prefix caching and mixture-of-experts (MoE) routing. To address these problems, we propose AgentReplay, a configurable trace record-and-replay framework for agent serving. AgentReplay records input/output tokens, expert selections, request dependencies, and tool durations. During replay, it forces the recorded output tokens while performing normal autoregressive computation, with optional controls for MoE expert selection and tool delays. This allows different systems to execute the same recorded workload while retaining their own batching, scheduling, and parallelization decisions. We further separate trajectory generation from performance evaluation, enabling compatible smaller models to replay long-horizon traces collected with more capable models. Our experiments show that token-wise replay in AgentReplay effectively eliminates workload variation that greedy decoding and length-wise replay cannot avoid, enabling fairer performance comparisons across serving configurations.

↑