arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SCX Router:结合解码器-KV分类器与真实任务本体的流式零样本模型选择

SCX Router: Streaming Zero-Shot Model Selection with a Decoder-KV Classifier and a Real-World Task Ontology

Ihor Stepanov, Aleksandr Smechov, Mykhailo Shtopko, Dmytro Vodianytskyi, Oleksandr Lukashov

arXiv 2609.02292首次发表:更新:

发表机构

Knowledgator; SCX.ai Holdings Limited(诺莱格特; SCX.ai控股有限公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出基于GLiClass的轻量级SCX Router,结合Qwen3解码器与任务本体实现流式零样本模型选择,在LiveBench子集上性能优于候选模型均值,1000任务子集累计Top-1得分达0.707,优于最强固定模型。

AI 中文摘要

大型语言模型(LLMs)的快速普及及其应用场景的日益多样化,带来了独特的优化机遇:针对每项任务选择合适的模型,同时在任务层面优化速度、成本与质量。然而,推理端点在质量、价格、延迟、上下文支持、工具使用、领域专长及推理行为等方面差异极大,这种异质性使得手动启发式方法难以维护,且仅靠手动方法不太可能始终实现理想的速度-成本-质量权衡。我们提出了SCX Router(简称\router{}),这是一种基于GLiClass的轻量级路由系统,无需自回归生成即可为每个推理时模型标签分配适配度评分。发布的0.6B参数检查点结合了Qwen3解码器与浅层双向评分器,其解码器-KV执行路径在会话中保留纯文本键值缓存,仅编码新对话轮次,并评估临时候选标签令牌且不将其添加到持久缓存中。该检查点还可预测任务类型、难度、推理模式及预期输出长度,并支持自定义零样本标签。为生成任务,我们构建了包含23个任务族、115种任务类型、345种可路由子类型、1173个合成示例及30个正交领域轴的任务本体。基于此结构,我们生成了150000个经验证器评分的任务与15000个开放式任务,随后在这些任务上训练Qwen3解码器,同时明确区分学习到的请求预测与任务级策略,包括适配性、成本、缓存复用、安全性及主权等方面。在6个LiveBench子集上,该路由系统的性能优于候选模型的平均值;在选定的1000任务子集上,其累计Top-1得分为0.707,而最强固定模型的得分为0.696,且存在依赖于基准的提升。

英文摘要

The rapid proliferation of large language models (LLMs) and the growing diversity of their applications presents a unique optimization opportunity: selecting the right model for the task, while optimizing for speed, cost, and quality at a per-task level. However, inference endpoints can vary widely in quality, price, latency, context support, tool use, domain expertise, and reasoning behavior. This heterogeneity makes manual heuristics difficult to maintain and unlikely to achieve consistently favorable speed--cost--quality trade-offs on their own. We introduce \router{}, a lightweight GLiClass-based router that assigns a suitability score to each inference-time model label without autoregressive generation. The released 0.6B-parameter checkpoint combines a Qwen3 decoder with a shallow bidirectional scorer. Its decoder-KV execution path preserves a text-only key--value cache across a session, encodes only new dialogue turns, and evaluates transient candidate-label tokens without adding them to the persistent cache. The same checkpoint also predicts task type, difficulty, reasoning mode, and expected output length, and supports custom zero-shot labels. For task generation, we construct a task ontology with 23 families, 115 task types, 345 routable subtypes, 1,173 synthetic examples, and an orthogonal axis of 30 domains. Using this structure, we generate 150,000 verifier-scored tasks and 15,000 open-ended tasks. We then train the Qwen3 decoder on these tasks, while explicitly separating learned request prediction from per-task policies for attributes such as eligibility, cost, cache reuse, safety, and sovereignty. Across six LiveBench subsets, the router outperforms the mean candidate; on the selected 1,000-task subset, it achieves an aggregate top-1 score of 0.707 versus 0.696 for the strongest fixed model, with benchmark-dependent gains.

Comments20 pages, 10 tables, 6 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑