arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.21621cs.LG

FrugalSOT——面向模型的节俭搜索

FrugalSOT - Frugal Search Over the Models

  • Vellore Institute of Technology(韦洛尔理工学院)

机构由 AI 辅助整理,请以论文原文为准。

Pradheep P, Yuvanesh S, Harish KB, Keerthan Saai Reddy S, Joshva Devadas T, Naveenkumar J, Hemalatha K

AI总结:

FrugalSOT是面向设备端NLP的资源感知模型选择架构,通过自适应阈值匹配模型复杂度,在树莓派5上可显著降低推理时间与资源消耗,同时保障输出相关性。

AI中文摘要:

在设备端自然语言处理(NLP)任务中,树莓派5(Raspberry Pi 5)等嵌入式硬件的有限资源要求采用高效推理策略。本文提出FrugalSOT(Frugal Search Over The Models,面向模型的节俭搜索),一种用于设备端NLP推理的资源感知型模型选择架构。FrugalSOT通过提取提示词长度、命名实体密度、句法复杂度等特征,估算每个请求的复杂度。请求首先被发送至最简易的模型,该模型需满足相关性阈值要求;若该模型输出未达阈值,则将请求发送至更复杂的模型。需注意,相关性阈值会在后台通过低通滤波机制,利用过往验证结果进行自适应更新,从而适配不断变化的输入模式。在树莓派5上开展的实验结果显示,与单模型基线方法相比,FrugalSOT可显著降低平均推理时间和整体计算资源消耗,同时不会像最复杂模型那样过度牺牲输出相关性。这些结果证实,自适应模型选择可在资源有限的设备上实现高效、高质量的自然语言处理推理。

英文摘要:

In on-device NLP tasks, limited resources of embedded hardware, such as the Raspberry Pi 5, require efficient inference strategies. This paper introduces FrugalSOT (Frugal Search Over The Models), a resource-aware model selection architecture for on-device NLP inference. FrugalSOT estimates each request's complexity by extracting features such as prompt length, named entity density, and syntactic complexity. The request is first made to the least complex model that is likely to pass a relevance threshold. If the output of that model falls short of the threshold, the request is made to a more complex model. It is important to note that the relevance threshold undergoes continuous updates in the background. using past validation outcomes in an adaptation process using a low-pass filtering mechanism, thus imparting adaptation to changing input patterns. Experimental results achieved on a Raspberry Pi 5 show that FrugalSOT reduces average inference time and overall computational resource use to a significant extent compared to a single-model baseline approach, without compromising output relevance to the same extent as the most sophisticated model. These results confirm that adaptive model selection can enable efficient, high-quality natural language processing inference on limited devices.

补充信息

↑