发表机构
AI lab of enakronIC PC; Electrical & Computer Engineering University of Patras(EnakronIC公司人工智能实验室; 帕特雷大学电气与计算机工程学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对本地部署检索增强型工厂智能体适配硬件难的问题,提出测量驱动的子网络选择方法,经案例研究验证可挽回模型质量损失并适配多异构边缘层级。
AI 中文摘要
本地部署的助手可为工厂工人提供对机器文档的对话式访问,但能完成该任务的模型很少能适配车间硬件。研究表明,在结构压缩和基于检索的适配后,模型规模不再是适配后答案质量的可靠预测指标:通用能力随参数数量几乎线性下降,而经评估的检索增强答案质量则不然。因此,将部署视为适配后的选择问题,在可配置的通用能力下限和内存预算下,根据评估的答案质量和实测的设备吞吐量,为每个设备分配一个子网络;仅优化规模、速度或质量的规则会在能力或吞吐量上做出妥协。采用三明治式原位蒸馏训练的权重共享超网络使该选择成本低廉。在一份制造手册案例研究中,提取操作的成本为未剪枝模型评估质量的13.7%,而基于检索的蒸馏将其拉回至4.6%以内,挽回了三分之二的损失,且同一助手可在三个异构边缘层级以1.3至5瓦的待机功率运行。
英文摘要
On-premise assistants can give factory workers conversational access to machine documentation, but models capable of the task rarely fit shop-floor hardware. We show that after structural compression and retrieval-grounded adaptation, model size is no longer a reliable predictor of adapted answer quality: general capability falls almost linearly with parameter count, while judged retrieval-augmented answer quality does not. We therefore treat deployment as a post-adaptation selection problem, committing one sub-network per device on judged answer quality and measured on-device throughput under a configurable general-capability floor and memory budget; rules that optimize size, speed, or quality alone each give up capability or throughput. A weight-shared supernetwork trained with sandwich-style in-place distillation keeps this selection inexpensive. In a manufacturing-manual case study, extraction costs 13.7 percent of the unpruned model's judged quality and retrieval-grounded distillation returns it to within 4.6 percent, recovering two thirds of the loss, and the same assistant runs across three heterogeneous edge tiers at 1.3 to 5 watts standby.