AI 中文总结
MergeHEIR提出合并后适应框架,通过零空间投影缓解模型合并导致的幻觉增加,同时保留继承的专家能力,在24组对比中实现更优的幻觉-保留权衡。
AI 中文摘要
模型合并将任务专用专家模型整合为单一可部署模型。然而,我们表明这种能力整合会带来幻觉增加这一合并税:在8种模型合并方法中,每个合并后的检查点都比其组成专家的平均幻觉率更高。一种直观的方法是使现有的幻觉缓解方法适应合并后的模型,但这种无约束的适应会破坏继承的能力,在幻觉缓解与专业知识保留之间造成矛盾。为应对这一挑战,我们引入MergeHEIR,这是一种合并后适应框架,旨在减少这种合并税,同时保留从初始专家继承的专业知识。使用小型专家任务校准集,MergeHEIR通过SVD从合并检查点收集的任务特定激活构建逐层零空间投影器,并定期将累积的合并后位移投影到所得零空间上,以保留继承的专业知识。理论上,我们建立了最小失真和最大维度保证,刻画了阈值控制的适应-保留权衡,并将扰动保证扩展到有限校准数据之外。在跨越3种多模态大语言模型配置和8种模型合并方法的24组成对比较中,MergeHEIR持续缓解幻觉,同时基本保留继承的专业知识,展示了更优的幻觉-保留权衡。
英文摘要
Model merging consolidates task-specialized experts into a single deployable model. However, we show that such capability consolidation incurs a merging tax of increased hallucination: across 8 model-merging methods, every merged checkpoint exhibits a higher hallucination rate than the average of its constituent experts. An intuitive approach is to adapt existing hallucination-mitigation methods to the post-merge model, yet this unconstrained adaptation disrupts inherited capabilities, creating a tension between hallucination mitigation and expertise retention. To tackle this challenge, we introduce MergeHEIR, a post-merge adaptation framework designed to reduce this merging tax while preserving expertise inherited from initial experts. Using small expert-task calibration sets, MergeHEIR constructs layer-wise null-space projectors via SVD from task-specific activations collected from the merged checkpoint, and periodically projects the accumulated post-merge displacement onto the resulting null spaces to preserve inherited expertise. Theoretically, we establish minimum-distortion and maximum-dimensionality guarantees, characterize the threshold-controlled adaptation-retention trade-off, and extend perturbation guarantees beyond finite calibration data. Across 24 paired comparisons spanning three MLLM configurations and 8 model-merging methods, MergeHEIR consistently mitigates hallucination while largely preserving inherited expertise, demonstrating a more favorable hallucination-retention trade-off.