arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向设备-边缘多模态推理的基于残差向量量化的任务导向视觉特征压缩

Task-Oriented Visual Feature Compression via Residual Vector Quantization for Device-Edge Multimodal Inference

Luning Pang, Cheng Yuan, Jiawei Shao, Mingtao Huang, Yuan Shen

arXiv 2609.37090首次发表:更新:

发表机构

Tsinghua University; Beijing National Research Center for Information Science and Technology; China Telecom; Institute of Artificial Intelligence (TeleAI), China Telecom(清华大学; 北京信息科学与技术国家研究中心; 中国电信; 中国电信人工智能研究院(TeleAI))

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对设备-边缘多模态推理中视觉特征传输延迟高的问题,提出查询引导的任务导向特征压缩方法Q-TOFC,利用残差向量量化和查询相关聚合降低负载53.6%,同时保持任务性能。

AI 中文摘要

大型多模态模型(LMMs)支持多种视觉理解和推理任务,但在资源受限的设备上完全运行这些模型往往不切实际。设备-边缘协同推理减少了设备计算负担,然而在带宽受限的上行链路上传输视觉数据可能引入大量延迟。任务导向特征压缩(TOFC)通过特征聚合和熵编码减少了传输负载。然而,连续特征编码仍然代价高昂,且与查询无关的聚合可能丢弃任务相关的局部证据。我们提出了查询引导的任务导向特征压缩(Q-TOFC)用于设备-边缘多模态推理。Q-TOFC采用残差向量量化(RVQ)将每个融合特征编码为紧凑的码本索引序列,降低了其表示成本,并允许传输更多特征。它进一步将查询相关性纳入特征聚合,并使用量化误差补偿适配器来减轻离散量化引入的失真。在七个多模态基准上的实验表明,与TOFC相比,Q-TOFC将视觉负载减少了53.6%,同时保持了相当的平均归一化任务性能。端到端延迟评估进一步证明了在带宽受限上行链路下的更低延迟。

英文摘要

Large multimodal models (LMMs) support diverse visual understanding and reasoning tasks but are often impractical to run entirely on resource-constrained devices. Device-edge co-inference reduces device computation, yet transmitting visual data over bandwidth-limited uplinks can introduce substantial delay. Task-oriented feature compression (TOFC) reduces the payload through feature aggregation and entropy coding. However, continuous-feature coding remains costly, and query-agnostic aggregation may discard task-relevant local evidence. We propose query-guided task-oriented feature compression (Q-TOFC) for device-edge multimodal inference. Q-TOFC employs residual vector quantization (RVQ) to encode each merged feature as a compact sequence of codebook indices, reducing its representation cost and allowing more features to be transmitted. It further incorporates query relevance into feature aggregation and uses a quantization error compensation adapter to mitigate the distortion introduced by discrete quantization. Experiments on seven multimodal benchmarks show that Q-TOFC reduces the visual payload by 53.6% relative to TOFC while maintaining comparable average normalized task performance. End-to-end latency evaluations further demonstrate lower latency under bandwidth-constrained uplinks.

Comments13 pages. Submitted to IEEE Transactions on Mobile Computing

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑