arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SCOPE-Router:面向执行导向任务的成本感知开放集视觉语言模型路由

SCOPE-Router: Cost-Aware Open-Set VLM Routing for Execution-Oriented Tasks

Tao Yu, Yifei Qu, Zhiqing Cui, Pengfei Zhou, Zhongtian Luo, Yujia Yang, Shenghua Chai, Haopeng Jin, Zhenghao Zhang, Xinming Wang, Hongzhu Yi, Wangbo Zhao, Zhenglin Wan, Yan Huang, Yeshani, Jinwen Luo, Yang You

arXiv 2608.12127首次发表:更新:

发表机构

CASIA; UCAS; NUS; Tencent(中国科学院自动化研究所; 中国科学院大学; 新加坡国立大学; 腾讯)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出SCOPE-Router与CRM+RCCR,构建执行导向VLM路由基准VLM-ExecRouterBench,解决现有VLM路由局限,在多基准上取得更优性能。

AI 中文摘要

模型路由旨在为每个查询从候选模型池中选择最合适的模型,以平衡性能与成本。现有视觉语言模型(VLM)路由研究局限于传统视觉问答(VQA)评估,缺乏针对开放集场景的系统性校准优化,且采用的训练目标通过softmax归一化稀释了多正样本信号,未纳入成本因素。针对这些局限,本文作出三项贡献:(1)VLM-ExecRouterBench,首个覆盖代码、智能体(Agentic)、搜索领域的执行导向VLM路由基准,包含11个候选模型,价格跨度近两个数量级;(2)SCOPE-Router,一种双塔路由器,通过混合校准(随机/诊断/多样性采样)构建模型行为轮廓,使新模型无需重新训练即可加入路由;(3)CRM+RCCR,一种架构无关的成本感知目标,通过逐对独立评分将成本偏好编码为连续相关性目标,消除多正样本稀释,同时将具有相似路由偏好的查询在路由空间中正则化至更接近。实验表明,SCOPE-Router在所有三个基准上取得最佳排名分数,在分布外(OOD)设置下较亚军高出1.84分,在双重分布外开放集评估下高出6.75分;将CRM+RCCR应用于四种不同路由器时,排名分数提升1.25至6.21分。

英文摘要

Model routing aims to select the most suitable model from a candidate pool for each query, balancing quality and cost. Existing VLM routing research is limited to traditional VQA evaluation, lacks systematic calibration optimization for open-set scenarios, and employs training objectives that dilute multi-positive signals via softmax normalization without incorporating cost. We address these limitations with three contributions: (1)VLM-ExecRouterBench, the first execution-oriented VLM routing benchmark covering Code, Agentic, and Search domains with 11 candidate models spanning nearly two orders of magnitude in pricing; (2)SCOPE-Router, a dual-tower router that matches queries to model behavior profiles constructed via hybrid calibration (random/diagnostic/diversity sampling), enabling new models to join routing without retraining; (3)CRM+RCCR, an architecture-agnostic cost-aware objective that encodes cost preference into continuous relevance targets through per-pair independent scoring, eliminating multi-positive dilution while regularizing queries with similar routing preferences to be closer in the routing space. Empirically, SCOPE-Router achieves the best Rank Score on all three benchmarks, surpassing the runner-up by 1.84 points under OOD settings and by 6.75 points under doubly OOD open-set evaluation. When applied to four diverse routers, CRM+RCCR improves Rank Score by 1.25--6.21 points.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑