发表机构
KAIST; Sony Group Corporation(韩国科学技术院; 索尼集团公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文针对现有全模态嵌入方法的结构分离不足问题,提出Syn-Omni框架,通过OME-LoRA与PSR实现结构化专业化和渐进式跨模态协作,在81项多模态任务上性能优于基线。
AI 中文摘要
全模态嵌入自然包含跨异构输入的共享表示和模态特定特征,但现有全模态嵌入方法常依赖混合模态数据上的单一共享参数空间,限制了通用表示与模态特定表示的结构分离。为解决该问题,本文提出Syn-Omni,这是一个兼具模态专业化与受控跨模态协作的结构化全模态适配统一框架。具体而言,我们引入正交模态专家LoRA(OME-LoRA),将适配分解为用于通用语义的共享LoRA路径和用于模态感知专业化的模态专家LoRA路径;此外,渐进协同路由(PSR)使专家先建立模态特定先验,再逐步与其他模态专家交互以实现跨模态协同。在涵盖图像、视频、音频及视听模态的81项多样化任务上评估,Syn-Omni始终优于全模态基线,证明了结构化专业化与跨模态渐进式协作的有效性。
英文摘要
Omnimodal embeddings naturally involve both shared representations and modality-specific features across heterogeneous inputs. However, existing omnimodal embedding methods often rely on a single shared parameter space over mixed-modality data, limiting structural separation between universal and modality-specific representations. To address this, we propose Syn-Omni, a unified framework for structured omnimodal adaptation with modality specialization and controlled cross-modal collaboration. Specifically, we introduce Orthogonal Modality-Expert LoRA (OME-LoRA), which decomposes adaptation into a shared LoRA path for universal semantics and modality-expert LoRA paths for modality-aware specialization. Furthermore, Progressive Synergy Routing (PSR) enables experts to first establish modality-specific priors, then gradually interact with other modality-experts for cross-modal synergy. Evaluated across 81 diverse tasks spanning image, video, audio, and audiovisual modalities, Syn-Omni consistently outperforms omnimodal baselines, demonstrating the effectiveness of structured specialization and cross-modal progressive collaboration.
CommentsAccepted to EMNLP 2026 (Long, Findings). Code: https://github.com/sony/syn-omni