arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

潘多拉的AI模型路由盒:带代价价值估计的高效分配

Pandora's AI Model Routing Box: Efficient Allocation with Costly Value Estimation

Adam Fisch, Shubhendu Trivedi, Fantine Huot, William W. Cohen, Michael Kaisers, Mirella Lapata, Kate Larson, Jacob Eisenstein

arXiv 2608.20316首次发表:更新:

发表机构

Google DeepMind(谷歌DeepMind)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对异构AI系统的查询路由问题,提出带代价价值估计的Pandora's Router集中式策略及Pandora's Bidder去中心化策略,可在保证路由质量的同时减少高成本估计器的使用,还能在不同场景优化分配效率。

AI 中文摘要

由多个模型、架构、工具或推理时设置组成的异构AI系统,可通过将查询路由到能以最低成本最有效回答问题的专家,来提升质量与效率。路由需要估计每个专家的预期回报,但这种价值估计存在成本:廉价估计器(如基于嵌入的预测器)速度快但噪声大,而精确估计器(如可访问检索结果或部分推理轨迹的微调模型)成本高。我们将这种权衡形式化为经典的带代价检查的最优搜索问题——潘多拉的盒子问题。在高斯信号模型下,所得策略具有闭式的信息价值表达式,可针对每个专家和输入,判断细化价值估计是否值得其成本。我们将这种集中式策略称为Pandora's Router(潘多拉路由器),并将其扩展到去中心化场景,即Pandora's Bidder(潘多拉竞标者):专家在接受报价以认领查询前,可独立决定是否投资自我评估。在三个领域(标准多LLM基准、检索增强专家、具有可变推理时推理的LLM)的实验表明,Pandora's Router的路由质量与穷尽估计相当,但查询昂贵估计器的频率低得多;在去中心化场景中,当竞争估计准确时,信息价值推理可提升分配效率,然而当竞争估计噪声大时,它会以牺牲其他方为代价提升策略性专家的效用。

英文摘要

Heterogeneous AI systems composed of multiple models, architectures, harnesses, or inference-time settings can improve quality and efficiency by routing queries to the specialist who can answer most effectively at the lowest cost. Routing requires estimating each specialist's expected return, but this value estimation has a cost. Cheap estimators (e.g., embedding-based predictors) are fast but noisy, while accurate estimators (e.g., fine-tuned models with access to retrieval results or partial reasoning traces) are expensive. We formalize this tradeoff as an instance of Pandora's Box, the classical problem of optimal search with costly inspection. Under a Gaussian signal model, the resulting policies have closed-form value-of-information expressions that determine, for each specialist and input, whether refining the value estimate is worth its cost. We call the centralized policy Pandora's Router. We extend this to a decentralized setting, Pandora's Bidder, where specialists independently decide whether to invest in self-assessment before accepting an offered price to claim a query. Experiments across three domains---a standard multi-LLM benchmark, retrieval-augmented specialists, and LLMs with variable inference-time reasoning---show that Pandora's Router matches the routing quality of exhaustive estimation, while querying the expensive estimator far less often. In the decentralized setting, value-of-information reasoning improves allocative efficiency when competing estimates are accurate; when competing estimates are noisy, however, it can increase the strategic specialist's utility at the expense of others.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑