RouteGuard:当互补性不足时,对LLM多智能体系统中的路由增益进行认证
RouteGuard: Certifying Routing Gain in LLM Multi-Agent Systems When Complementarity Is Not Enough
浏览论文内容
中文总结 AI 辅助
针对LLM多智能体路由的部署问题,提出RouteGuard框架,通过分解增益、结合Le Cam下界等实现认证,在两个基准中验证了其有效性。
中文摘要 AI 辅助
多智能体LLM系统在由模型支持的顾问之间进行路由,但部署者在部署前很少知道路由是否真的有帮助。主流路由方法优化门控的AUC,并假设顾问的互补性足以保证路由有效。我们证明这两个假设都不能决定可部署的增益。我们引入了RouteGuard,一个部署认证框架。路由增益可分解为G = πΔ_E,可实现的增益由条件遗憾泛函Φ决定,而非AUC。有限样本认证区间带有匹配的Le Cam下界,在固定活动类上具有恒定尖锐度,且存在鲁棒性相变。在两个基准测试中,该框架起到了护栏作用:在RouterBench(11个跨系列模型)上,结果取决于采样单元:在提示级采样下,协议认证了优于GPT-4的增益,而在工作负载集群重采样下则拒绝认证,因为增益取决于86个工作负载单元中的3个;在OpenRCA(三个Gemini顾问)上,顾问在统计上是冗余的:在我们测试的所有池(221个RouterBench池和三个OpenRCA分布)中,实现的神谕性能等于或低于独立性基线,因此协议正确拒绝认证。一个预先注册的半合成对照验证了校准:当m≥m*时,协议认证真正的增益,而不会认证真实的空情况。代码和冻结的工件将在发布版本中发布。
英文摘要
Multi-agent LLM systems route among model-backed advisors, yet a deployer rarely knows before shipping whether routing will help at all. Prevailing routers optimize a gate's AUC and presume that advisor complementarity suffices. We show that neither determines the deployable gain. We introduce RouteGuard, a deployment-certification framework. Routing gain decomposes as $G = πΔ_E$, and the achievable gain is governed by a conditional-regret functional $Φ$, not by AUC. A finite-sample certification bracket comes with a matching Le Cam lower bound, constant-sharp over the fixed-activity class, and a robustness phase transition. On two benchmarks the framework acts as a guardrail. On RouterBench (11 cross-family models) the verdict depends on the sampling unit: the protocol certifies a gain over GPT-4 under prompt-level sampling and withholds it under workload-cluster resampling, because the gain rests on 3 of 86 workload cells. On OpenRCA (three Gemini advisors) the advisors are statistically redundant: the realized oracle sits at or below the independence baseline in all pools we tested (221 RouterBench pools and three OpenRCA distributions), so the protocol correctly refuses to certify. A pre-registered semi-synthetic control confirms calibration: the protocol certifies a genuine gain once $m \ge m^\star$ and does not certify a true null. Code and frozen artifacts will be released with the published version.
发表机构
- University of Miami(迈阿密大学)
- Google(谷歌公司)
机构由 AI 辅助整理,请以论文原文为准。