arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2512.08524cs.CVcs.CL

超越真实权重:用于稳定量化 的超复数表示

Beyond Real Weights: Hypercomplex Representations for Stable Quantization

  • Artificial Intelligence Department, RobotBulls Labs(机器人bulls实验室人工智能部门)
  • Machine Intelligence Lab (MILab), North South University(北南大学机器智能实验室)

机构由 AI 辅助整理,请以论文原文为准。

Jawad Ibn Ahad, Maisha Rahman, Amrijit Biswas, Muhammad Rafsan Kabir, Robin Krambroeckers, Sifat Momen, Nabeel Mohammed, Shafin Rahman

更新

AI总结:

本文提出了一种基于超复数乘法的渐进式重新参数化策略,用于压缩多模态语言模型,实现参数和计算量的显著减少,同时保持模型性能。

AI中文摘要:

多模态语言模型(MLLMs)需要大量的参数容量来对齐高维视觉特征与语言表示,这使得它们计算量大,难以高效部署。我们介绍了一种渐进重新参数化策略,通过逐步将密集的前馈网络块替换为紧凑的参数化超复数乘法(PHM)层来压缩这些模型。残差插值计划,结合轻量级重建和知识蒸馏损失,确保PHM模块在训练过程中继承其密集对应物的功能行为。这种转换在保持强多模态对齐的同时,实现了显著的参数和FLOP减少,从而在不降低输出质量的情况下实现更快的推理。我们在多个视觉-语言模型(VLMs)上评估了该方法。我们的方法在保持与基模型相当的性能的同时,实现了模型大小和推理延迟的显著减少。渐进式的PHM替换因此为更高效的多模态推理提供了一条架构兼容的路径,并补充了现有的低比特量化技术。

英文摘要:

Multimodal language models (MLLMs) require large parameter capacity to align high-dimensional visual features with linguistic representations, making them computationally heavy and difficult to deploy efficiently. We introduce a progressive reparameterization strategy that compresses these models by gradually replacing dense feed-forward network blocks with compact Parameterized Hypercomplex Multiplication (PHM) layers. A residual interpolation schedule, together with lightweight reconstruction and knowledge distillation losses, ensures that the PHM modules inherit the functional behavior of their dense counterparts during training. This transition yields substantial parameter and FLOP reductions while preserving strong multimodal alignment, enabling faster inference without degrading output quality. We evaluate the approach on multiple vision-language models (VLMs). Our method maintains performance comparable to the base models while delivering significant reductions in model size and inference latency. Progressive PHM substitution thus offers an architecture-compatible path toward more efficient multimodal reasoning and complements existing low-bit quantization techniques.

补充信息

↑