HACO:用于可靠大语言模型系统的套期保值代理计算
HACO: Hedged Agent Computing for Reliable LLM Systems
浏览论文内容
中文总结 AI 辅助
研究大语言模型代理在长期工作流中角色到实例绑定边界的故障问题,提出HACO运行时控制方案,将角色请求视为可靠性约束选择问题,通过结合乐观排序与保守可靠性积累及经验收集更新配置文件,提高了模型在变化条件下的性能并降低成本。
中文摘要 AI 辅助
随着大语言模型(LLM)代理从孤立提示转向长期工作流,在角色到实例绑定边界处故障日益增多,当前服务、网络和查询条件下任务特定角色请求需分配给具体代理实例。现有代理系统研究虽改善了角色专业化等,但常假定固定稳定执行环境,限制了部署可靠性。本文提出套期保值代理计算(HACO),一种运行时控制方案,将每个角色请求视为候选代理实例上的可靠性约束选择问题,不同调用会自适应选择候选套期保值集,其分配规则结合乐观排序和保守可靠性积累,通过经验收集更新候选和链接配置文件。实验表明,HACO在变化部署条件下提高了鲁棒性和输出质量,且使用的令牌和延迟成本低于穷举并行执行。
英文摘要
As large language model (LLM) agents move from isolated prompting to longhorizon workflows, failures increasingly arise at the role-to-instance binding boundary, where task-specific role requests must be assigned to concrete agent instances under current service, network, and query conditions. Existing agent system research has improved role specialization, workflow topology, memory, and tool use, but often assumes a fixed stable execution environment. This assumption limits deployed reliability, because the same role request can exhibit different latency, failure probability, and output quality across agent instances operating under different service regions and network conditions. We propose Hedged Agent Computing (HACO), a runtime control scheme that treats each role request as a reliability-constrained selection problem over candidate agent instances, each coupling a role type, an LLM, and a concrete execution environment. Different from routing, HACO adaptively selects a hedge set of candidates for each invocation. Its allocation rule combines optimistic ranking, which prioritizes candidates with high estimated quality, reliability, and informative uncertainty, with conservative reliability accumulation, which stops selection only after the hedge set reaches a target success probability. Through experience harvesting, HACO updates candidate and link profiles from all executed candidate traces, including quality, success, latency, and network statistics. Experiments on various benchmarks, together with runtime degradation studies, show that HACO improves robustness and output quality under changing deployment conditions, while using lower token and latency cost than exhaustive parallel execution.