AI 中文总结
针对智能体多模态大语言模型合并时的弱任务退化与关键行为遗忘问题,提出无需训练的AgentPatch框架,通过由粗到细修复实现能力平衡,提升合并模型性能。
AI 中文摘要
智能体多模态大语言模型(MLLM)通过规划、工具使用及动态环境交互拓展了多模态感知与推理能力。然而当前模型针对特定工具或环境进行了专门化设计,难以整合为单一通用模型。本文提出智能体MLLM合并问题,并识别出两大挑战:一是非对称能力保留,即不同交互复杂度的能力保留不均,合并后产生弱任务;二是行为关键遗忘,即丢失关键动作可能破坏长时序执行。我们提出AgentPatch,一种无需训练的由粗到细修复框架:它选择稳定的合并主干网络,通过弱任务独特残差恢复(Weak-Task Unique Residual Recovery)恢复被稀释的弱任务特定信号,并应用智能体引导的行为关键补丁(Agent-Guided Behavior-Critical Patch),在明确的能力保护下恢复关键行为。AgentPatch生成单一静态检查点,无需路由或集成。在6个智能体与多模态基准上的实验显示,AgentPatch可提升多种合并主干网络,缓解弱任务性能下降,更好平衡弱任务恢复与互补搜索及智能体视觉处理能力的保留。代码可在this https URL获取。
英文摘要
Agentic multimodal large language models (MLLMs) extend multimodal perception and reasoning with planning, tool use, and interaction in dynamic environments. Yet current models are specialized for particular tools or environments, complicating consolidation into a single generalist. We formulate Agentic MLLM Merging and identify two challenges: asymmetric capability preservation, whereby capabilities with different interaction complexity are retained unevenly, producing weak tasks after merging, and behavior-critical forgetting, whereby losing decisive actions can derail long-horizon execution. We propose AgentPatch, a training-free coarse-to-fine repair framework. It selects a stable merged backbone, restores diluted weak-task-specific signals through Weak-Task Unique Residual Recovery, and applies an Agent-Guided Behavior-Critical Patch that recovers decisive behaviors under explicit capability protection. AgentPatch produces a single static checkpoint without routing or ensembles. Experiments across six agentic and multimodal benchmarks show that AgentPatch improves diverse merged backbones, alleviates weak-task degradation, and better balances weak-task recovery with the preservation of complementary search and agentic visual processing capabilities. Code is available at https://github.com/ziboshao/AgentPatch.