发表机构
Beijing Institute of Technology(北京理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出OnPTQ在线策略后训练量化框架,通过决策-后果风险评估关键状态,在多模态大语言模型上提升性能并减少正确性翻转。
AI 中文摘要
后训练量化(PTQ)降低了多模态大语言模型的部署成本,但校准通常以局部目标重建固定序列。这忽略了自回归反馈:量化引起的词元变化会重定向前缀并改变未来状态。然而,仅靠在线策略覆盖是不够的,因为许多决策不匹配几乎不影响未来的生成。我们提出OnPTQ,一种在线策略框架,在当前量化策略访问的轨迹上进行校准。在共享前缀上,OnPTQ识别量化侵蚀的边界,通过短的反事实回滚评估竞争词元,并将当前差异与分支后果结合为决策-后果风险。该风险优先考虑关键状态,而上下文锚定和轨迹刷新保持多模态行为并使校准与更新后的策略保持一致。我们进一步推导出决策-后果界限,将行为偏差与当前策略差异和动作条件下的未来值跨度联系起来。在多种低位设置下的视觉-语言和全模态Qwen模型中,OnPTQ提高了下游性能,并相对于相应的Dense/FP16参考产生更少的正确性翻转,且不改变部署的推理图。
英文摘要
Post-training quantization (PTQ) lowers deployment cost for multimodal large language models, but calibration typically reconstructs fixed sequences with local objectives. This overlooks autoregressive feedback: a quantization-induced token change redirects the prefix and changes future states. Yet on-policy coverage alone is insufficient because many decision mismatches barely affect future generation. We propose OnPTQ, an on-policy framework that calibrates on trajectories visited by the current quantized policy. On shared prefixes, OnPTQ identifies quantization-eroded boundaries, evaluates competing tokens through short counterfactual rollouts, and combines current discrepancy with branch consequence into a Decision--Consequence risk. The risk prioritizes critical states, while context anchoring and trajectory refresh preserve multimodal behavior and keep calibration aligned with the updated policy. We further derive a Decision--Consequence bound linking behavioral deviation to current policy discrepancy and action-conditioned future-value span. Across vision--language and omni-modal Qwen models under multiple low-bit settings, OnPTQ improves downstream performance and yields fewer correctness flips against the corresponding Dense/FP16 references, without changing the deployed inference graph.
Commentspreprint