选择即检索,弃权(不执行)则不然:覆盖70个韩英动作的设备端工具路由
Selection Is Retrieval, Abstention Is Not: On-Device Tool Routing over 70 Korean-English Actions
浏览论文内容
中文总结 AI 辅助
针对设备端工具路由,研究发现选择可依赖BM25检索,而弃权(不执行)判断需神经编码器,神经排序器因延迟和内存受限而非准确性被弃用。
中文摘要 AI 辅助
一个调用工具的AI助手在每个请求上需要做出两个决策:调用哪个工具,以及是否有任何可用工具适用。在通常的设计中,单个语言模型通过发出调用或拒绝发出调用来同时做出这两个决策。在必须无服务器应答的设备上,语言模型正是使该设计代价高昂的原因,它在延迟和内存方面都主导了路由器。常见的替代方案是完全移除模型,转而使用检索器对本地动作目录进行排序。这种替代在这两个决策上并不对称。检索器对每个输入都返回其得分最高的候选,无法表明目录中不存在有效动作。我们早前的研究发现,将解码器约束到工具语法可以修复格式错误的输出,但并未改善选择。这种替代在每个决策上的代价尚未被衡量。我们分别评估了这两个决策,涉及600个韩语和英语请求以及一个包含70个本地动作的目录。路由器还可以请求缺失的槽位、回复或委派。目录内请求中有一半复用目录词汇,另一半则进行释义,从而将词汇重叠与所请求的动作分开。字符3-gram BM25在164个词汇匹配请求中选中了162个,在166个释义中选中了85个。将候选集限制为七个后,五次试验中释义的平均值提高到0.825。没有任何基于其得分特征的分类器能在目录内与目录外之间达到超过0.697的曲线下面积,而冻结的编码器multilingual-e5-base达到了0.806。仅使用该编码器进行弃权(不执行)判断,可将376个请求保留在本地,并在150个需要委派的请求中错误路由9个。弃权(不执行),而非选择,才是需要神经组件的地方。神经排序器改善了所有质量指标,但其被拒绝是出于延迟和内存原因,而非准确性。
英文摘要
An AI assistant that calls tools makes two decisions on every request: which tool to invoke, and whether any available tool applies. In the usual design a single language model makes both, by emitting a call or by declining to emit one. On a device that has to answer without a server, the language model is what makes that design expensive, dominating both the latency and the memory of the router. The common alternative is to remove the model completely and rank the catalog of local actions with a retriever instead. That substitution is not symmetric across the two decisions. A retriever returns its highest-scoring candidate for every input and cannot signal that the catalog holds no valid action. Our earlier study found that constraining a decoder to a tool grammar repairs malformed output without improving the choice. What the substitution costs in each decision has not been measured. We evaluate the two decisions separately over 600 Korean and English requests and a catalog of 70 local actions. The router may also ask for a missing slot, reply, or delegate. Half the in-catalog requests reuse catalog vocabulary and half paraphrase it, separating lexical overlap from the action requested. Character 3-gram BM25 selects 162 of 164 lexically matched requests and 85 of 166 paraphrases. Restricting the candidate set to seven raises the paraphrase figure to a mean of 0.825 over five trials. No classifier over its score features separates in-catalog from out-of-catalog above 0.697 area under the curve, where the frozen encoder multilingual-e5-base reaches 0.806. Using that encoder for abstention alone keeps 376 of the requests local and misroutes 9 of the 150 needing delegation. Abstention, not selection, is where a neural component is required. A neural ranker improves every quality metric and is rejected on latency and memory rather than accuracy.
发表机构
- Redrob
机构由 AI 辅助整理,请以论文原文为准。