From Inheritance to Saturation: Disentangling the Evolution of Visual Redundancy for Architecture-Aware MLLM Inference Acceleration
从继承到饱和:解构视觉冗余的演化以实现架构感知的MLLM推理加速
机构 * University of Science and Technology of China(中国科学技术大学) ; Wuhan University(武汉大学) ; Ping An Technology (Shenzhen) Co., Ltd.(平安科技(深圳)有限公司)
AI总结 本文提出HalfV框架,通过统一剪枝策略缓解内在视觉冗余,并根据具体表现适应性处理二次饱和冗余,实现跨架构的高效推理。
Comments 16 pages, 14 figures, plus appendix, accepted at ACL 2026