arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

压缩大语言模型重编程用于车载网络中的视觉辅助波束预测

Compressed LLM Reprogramming for Vision-Aided Beam Prediction in Vehicular Networks

Kai Dong, Lei Wang, Changyi Li, Sergiy A. Vorobyov, Stefan Werner

arXiv 2609.35459首次发表:更新:

发表机构

Aalto University(阿尔托大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出LLMBP-Lite,通过三维剪枝和知识蒸馏压缩LLM重编程模型,实现65倍推理加速且保持精度,缓解车载网络中数据效率与部署效率的权衡。

AI 中文摘要

基于大语言模型(LLM)重编程的波束预测通过将预训练语言模型适配于车对基础设施(V2I)波束预测,展现出强大的数据效率,然而由此产生的模型复杂度使得此类方法在时延敏感部署中不切实际。我们提出了LLMBP-Lite,一个紧凑的LLM重编程波束预测框架。它通过沿三个互补维度进行剪枝来利用结构冗余:Transformer深度、源原型词汇表大小和提示长度。此外,可选地采用知识蒸馏以在压缩后保持预训练的表征能力。在真实世界数据集上的实验表明,LLMBP-Lite相比未压缩的基于LLM的基线实现了65倍的推理加速,同时保持预测准确性,并在有限训练数据下持续优于循环基线。这些结果表明,在车载网络中,数据效率与部署效率之间的权衡可以得到显著缓解。

英文摘要

Large Language Model (LLM) reprogramming-based beam prediction demonstrates strong data efficiency by adapting pretrained language models for vehicle-to-infrastructure (V2I) beam prediction, yet the resulting model complexity makes such approaches impractical for latency-sensitive deployment. We propose LLMBP-Lite, a compact LLM-reprogrammed framework for beam prediction. It leverages structural redundancy through pruning along three complementary dimensions: Transformer depth, source-prototype vocabulary size, and prompt length. Additionally, knowledge distillation can be optionally employed to maintain the pretrained representational capacity after compression. Experiments on the real-world dataset demonstrate that LLMBP-Lite achieves a $65\times$ inference speedup over the uncompressed LLM-based baseline while maintaining prediction accuracy and consistently outperforming recurrent baselines under limited training data. These results demonstrate that the tradeoff between data efficiency and deployment efficiency can be substantially mitigated in vehicular networks.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑