发表机构
National Yang Ming Chiao Tung University(国立阳明交通大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对大语言模型路由中忽视查询潜在难度问题,提出VDAR-Router框架,通过生成难度分析、检索相似示例来估计模型适用性,经实验验证该方法能实现更好成本性能权衡,明确查询分析有助于做出更可靠路由决策。
AI 中文摘要
大语言模型在实际系统中使用日益广泛,高效模型选择对降低部署成本至关重要。现有路由方法常从输入查询的表面语义或嵌入相似度估计模型适用性,可能忽略查询潜在难度。为此提出VDAR-Router,一种基于难度感知检索的路由框架。对每个输入查询先生成明确难度分析,再检索难度相似历史示例,基于检索记录估计候选模型适用性并通过奖励函数选模型。在三个数据集上实验表明,VDAR-Router比现有基线 consistently achieves better cost-performance trade-offs,案例研究显示明确查询分析有助于检索更相关示例并支持更可靠路由决策。
英文摘要
Large language models are increasingly used in practical systems, making efficient model selection important for reducing deployment cost. LLM routing has emerged as a practical solution for allocating each input query to an appropriate model under a desired cost-performance trade-off. Existing routing methods often estimate model suitability from the surface semantics or embedding similarity of the input query. However, such methods may ignore the underlying difficulty of a query, leading to suboptimal routing decisions. To address the challenge, we propose VDAR-Router, a difficulty-aware retrieval-based routing framework. For each input query, VDAR-Router first generates an explicit difficulty analysis. It then retrieves historical examples with similar difficulty profiles. Based on the retrieved records, it estimates candidate model suitability and selects the model using a reward function that considers both performance and cost. Experiments on three datasets show that VDAR-Router consistently achieves better cost-performance trade-offs than existing baselines. These results demonstrate the effectiveness of difficulty-aware retrieval for training-free LLM routing. Case studies further show that explicit query analysis helps retrieve more relevant examples and supports more reliable routing decisions.
CommentsAccepted by EMNLP 2026 Findings