AI 中文总结
研究多模型系统路由策略何时有意义,通过适应分层社会熵并引入基于扰动的度量来诊断失败模式,应用于EmbedLLM和RouterBench,发现HSE收益递减,KNN路由器扰动下鲁棒性差,提示路由稳定,揭示准确性与意义可能背离。
AI 中文摘要
多模型系统的路由策略几乎仅根据任务准确性和推理成本来评估。我们认为,与性能正交的两个属性决定了路由是否有意义。首先,参与者群体在行为上必须有差异:如果所有参与者反应相同,路由就毫无意义。其次,路由策略必须稳定:查询的表面形式变体应分配给同一个参与者。高任务准确性可能会违反这两个属性,因为路由器可能在冗余群体上运行或不一致地分配查询,无论性能如何都阻碍专业化。我们将分层社会熵(HSE)应用于语言模型群体,并引入基于扰动的鲁棒性度量来诊断这些失败模式。应用于EmbedLLM和RouterBench时,我们发现HSE呈现出强烈的收益递减,表明精心挑选的少于十个代理的子集就能在大群体中恢复大部分可用的多样性,这是一种实用的社会设计核心集启发式方法。我们还发现,KNN路由器从专业群体中获得准确性,但在扰动下鲁棒性会崩溃,而提示路由在所有扰动类型下都保持稳定,这说明准确性和意义可能会大幅背离。
英文摘要
Routing policies for multi-model systems are evaluated almost exclusively on task accuracy and inference cost. We argue that two properties, orthogonal to performance, determine whether routing is meaningful. First, the society of actors must be behaviourally differentiated: if all actors respond identically, routing is vacuous. Second, the routing policy must be stable: surface-form variants of a query should be assigned to the same actor. High task accuracy is compatible with violating both properties, since a router can operate over a redundant society or assign queries inconsistently, preventing specialisation regardless of performance. We adapt Hierarchic Social Entropy (HSE) to language-model societies and introduce a perturbation-based robustness metric to diagnose these failure modes. Applied to EmbedLLM and RouterBench, we find that HSE exhibits strong diminishing returns, suggesting that a curated subset of fewer than ten agents recovers most available diversity in a large pool -- a practical coreset heuristic for society design. We further find that KNN routers gain accuracy from specialist societies but collapse in robustness under perturbation, while prompted routing remains stable across all perturbation types -- illustrating that accuracy and meaningfulness can sharply diverge.