发表机构
Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳))
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出HEAR协议,连接智能体框架与推理引擎,实现工作流感知的高效LLM服务,在多个基准上显著提升速度且不损质量。
AI 中文摘要
LLM智能体日益执行涉及多轮推理、工具使用和并行智能体的复杂工作流。高效服务需要跨越两个具有互补信息的层的决策:智能体框架理解工作流依赖关系、上下文生命周期和执行目标,而推理引擎观察请求队列、KV缓存状态、资源压力和执行能力。现有接口未系统性地连接这些视图,限制了工作流感知的执行。HEAR,一种用于智能体LLM服务的双向框架-引擎配对协议。HEAR标准化了框架如何传达工作流意图和执行要求,以及引擎如何返回运行时状态、能力和结果。通过将协议语义与优化策略分离,HEAR支持多样化的协调策略,而无需改变工作流或模型语义。我们实例化HEAR用于在线缓存感知的运行时协调和面向智能体角色的工作负载感知的执行模式选择。在四个对话和研究智能体基准上,在内存受限、并发服务条件下,HEAR在SCBench上实现了1.61倍的批量加速,并将中位时间到首个令牌减少2.23倍。Mooncake表明,工作流意图和实时引擎状态在不同负载情况下提供互补的益处。在BrowseComp-Plus和DeepResearchBench上,工作负载特定的配置分别产生1.23倍和2.45倍的端到端加速,且未观察到任务质量下降。这些结果确立了HEAR作为高效智能体LLM服务的可复用协调基础。
英文摘要
LLM agents increasingly execute complex workflows involving multi-turn reasoning, tool use, and parallel agents. Efficient serving requires decisions that span two layers with complementary information: the agent harness understands workflow dependencies, context lifecycles, and execution objectives, whereas the inference engine observes request queues, KV-cache state, resource pressure, and execution capabilities. Existing interfaces do not systematically connect these views, limiting workflow-aware execution. HEAR, a bidirectional Harness--Engine Pairing protocol for agentic LLM serving. HEAR standardizes how the harness communicates workflow intent and execution requirements and how the engine returns runtime state, capabilities, and outcomes. By separating protocol semantics from optimization policies, HEAR supports diverse coordination strategies without changing workflow or model semantics. We instantiate HEAR for online cache-aware runtime coordination and workload-aware execution-mode selection for agent roles. Across four conversational and research-agent benchmarks under memory-constrained, concurrent serving, HEAR achieves a $1.61\times$ batch speedup and reduces median time-to-first-token by $2.23\times$ on SCBench. Mooncake shows that workflow intent and live engine state provide complementary benefits across load regimes. On BrowseComp-Plus and DeepResearchBench, workload-specific configurations yield $1.23\times$ and $2.45\times$ end-to-end speedups, respectively, without observed task-quality degradation. These results establish HEAR as a reusable coordination substrate for efficient agentic LLM serving.
CommentsJiaqi Zhao, Haodong Chen, and Jitai Hao contributed equally