arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

COVER:联盟路由的可识别评估

COVER: Identifiable Evaluation of Coalition Routing

Raghul Sugumar, Amrit Gopinath

arXiv 2608.28475首次发表:更新:

发表机构

Sri Sivasubramaniya Nadar College of Engineering(斯里·西瓦苏布拉马尼亚·纳达尔工程学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对多智能体系统路由效应无法通过端到端准确率识别的问题,提出可审计的COVER评估方法,在MuSiQue、HotpotQA等数据集及Llama模型上验证其可准确测量联盟路由遗憾,暴露选择空间且不人为制造路由胜利。

AI 中文摘要

当多智能体系统改变其团队时,它产生的消息和最终答案也会随之改变,因此端到端准确率差距本身并不能识别路由效应。我们引入了一种方法,即COVER,它是一种评估契约,在生成结果前会固定公共信息边界、下游栈G以及有限合法团队族。完全覆盖会识别出基于该栈的精确有限基准oracle遗憾。对于任意一组固定策略,执行它们各自选定团队的并集是每对策略对比的最小无假设支持,但不适用于绝对oracle遗憾。我们使用两个具有源ID不相交分割的受控表来测试该工具:在MuSiQue-12上,预先指定的特权阳性对照将遗憾从0.532改善至0.402;后续的公共接口对照达到0.424,而基准为0.554,但属于回顾性结果。在HotpotQA-4上,预先指定的公共直接评分器将遗憾从0.313改善至0.110。在固定栈Llama执行中,验证后的路由遗憾改善了0.190,而原始答案增益为0.010,其区间包含零。一个五族ToolSandbox变体转移验证对14个未触及的任务变体上的16个已声明团队进行了详尽评估(224/224有效行):已声明族oracle达到0.768的安全证据完成率,而前瞻性冻结路由器达到0.637(遗憾0.131),未达到预先声明的0.10标准。后续回顾性对照达到0.655,与所有工作者匹配,平均工作者数为4.57对5.00。因此,COVER可在不人为制造路由胜利的情况下暴露选择空间。交叉栈诊断显示绝对分数依赖于G,但未检测到路由器与最终确定器之间的可交互作用。COVER是一种可审计的测量方法,而非栈不变或通用智能体路由优越性的主张。

英文摘要

When a multi-agent system changes its team, it also changes the messages and final answer it produces, so an end-to-end accuracy gap does not by itself identify a routing effect. We introduce method, an evaluation contract that fixes a public information boundary, downstream stack G, and finite legal team family before outcomes are generated. Complete coverage identifies exact finite-benchmark oracle regret conditional on that stack. For any finite collection of frozen policies, executing the union of their distinct selected teams is the minimal assumption-free support for every pairwise policy contrast, though not for absolute oracle regret. Two controlled tables with source-ID-disjoint splits test the instrument. On MuSiQue-12, a pre-specified privileged positive control improves regret from 0.532 to 0.402; a later public-interface control reaches 0.424 versus 0.554 but is retrospective. On HotpotQA-4, a pre-specified public direct scorer improves regret from 0.313 to 0.110. In fixed-stack Llama execution, verified route regret improves by 0.190, while the raw-answer gain is 0.010 with an interval crossing zero. A five-family ToolSandbox variant-shift validation exhaustively evaluates 16 declared teams on 14 untouched task variants (224/224 valid rows): the declared-family oracle reaches 0.768 safe-evidence completion, while the prospectively frozen router gets 0.637 (regret 0.131), failing the predeclared 0.10 criterion. A later retrospective comparator reaches 0.655, matching all-workers with 4.57 versus 5.00 workers on average. Thus COVER exposes selection headroom without manufacturing a routing win. A crossed-stack diagnostic shows absolute scores depend on G but finds no detectable router-by-finalizer interaction. COVER is an auditable measurement methodology, not a claim of stack-invariant or universal agent-routing superiority.

Comments17 pages, 2 figures, 13 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑