设备优先反馈:面向移动端原生大语言模型驱动的神经网络架构搜索
Device-First Feedback: Toward Mobile-Native LLM-Driven Neural Architecture Search
- University of Würzburg(维尔茨堡大学)
- CAIDAS
- IFI
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该研究提出自动化移动端部署闭环流水线,针对CIFAR-10、CIFAR-100基准在三星平板验证,发现GPU微调无法保证移动端性能单调提升,需多数据集设备端测量。
AI中文摘要:
将大语言模型(LLM)生成的卷积神经网络部署到实际移动硬件上,仅靠GPU验证准确率是不够的:INT8格式的TensorFlow Lite导出、算子选择以及设备端延迟共同决定了模型是否可用。我们提出了一种自动化移动端部署流水线,形成从架构生成型LLM的QLoRA微调,到GPU评估、INT8导出、物理设备基准测试,再到训练语料库的门控增强的闭环。该流水线完全脚本化,可逐轮运行且无需人工干预,支持中断后恢复。我们在三星SM-P613平板(随机种子42,每轮20个模型,第0-6轮)上,针对CIFAR-10和CIFAR-100两个基准测试评估了相同的固定协议。在CIFAR-10上,第1轮通过门控,移动端部署得分较基线提升约25.6倍,量化准确率均值为46.9%;后续轮次提升了GPU准确率,但未通过非递减移动端门控。在CIFAR-100上,QLoRA前的基线保持最佳移动端得分;迭代轮次提升了GPU准确率(最高达26.2%),但无法超越第0轮的设备端表现,训练池在首次通过轮次后停滞在19个样本。两项研究共同表明,闭环GPU微调无法保证移动端性能的单调提升,尤其是在更复杂的分类任务中,需要多数据集、设备端测量来充分测试部署目标。我们发布了带95%置信区间的每轮指标、所有图表及完整复现命令。
英文摘要:
Deploying convolutional neural networks generated by large language models (LLMs) on real mobile hardware requires more than GPU validation accuracy: INT8 TensorFlow Lite export, delegate selection, and on-device latency jointly determine whether a model is usable. We present an automated mobile deployment pipeline that closes the loop from QLoRA fine-tuning of an architecture-generating LLM through GPU evaluation, INT8 export, and physical-device benchmarking to gated augmentation of the training corpus. The pipeline is fully scripted and runs cycle-by-cycle without manual intervention, with resume support after interruptions. We evaluate the same frozen protocol on two benchmarks, CIFAR-10 and CIFAR-100, on a Samsung SM-P613 tablet (seed 42, 20 models per cycle, cycles 0-6). On CIFAR-10, cycle 1 is gate-accepted and improves the mobile deployment score approximately 25.6x over the baseline with a mean quantized accuracy of 46.9%; later cycles raise GPU accuracy but fail the non-decreasing mobile gate. On CIFAR-100, the pre-QLoRA baseline retains the best mobile score; iterative rounds improve GPU accuracy (up to 26.2%) yet cannot surpass cycle 0 on-device, and the training pool stalls at 19 examples after the first accepted round. Together, the two studies show that closed-loop GPU fine-tuning does not guarantee monotonic mobile gains, especially on harder classification tasks, and that multi-dataset, on-device measurement is needed to stress-test deployment objectives. We release per-cycle metrics with 95% confidence intervals, all figures, and complete reproduction commands.