arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.06406cs.CL

分层Wasserstein合并用于多领域多任务学习:从专家到通才

Hierarchical Wasserstein Merging for Multi-Domain Multi-Task Learning: From Specialists to a Generalist

Ming Cheng, Jiaying Gong, Hoda Eldardiry

首次发表
浏览论文内容

中文总结 AI 辅助

提出分层Wasserstein合并(HWM)框架,通过建模表示分布并构建重心,实现多领域多任务学习中的专家聚合与通才训练,实验验证其有效性和泛化性。

中文摘要 AI 辅助

多领域多任务学习(MD-MTL)旨在构建一个在异构领域和任务上均表现良好的单一通才模型。然而,联合训练在分布偏移下常常遭受干扰。现有的模型合并方法大多作用于模型参数,而忽略了跨领域和任务间潜在表示分布的几何结构。为解决这些局限性,我们提出了分层Wasserstein合并(HWM),一种表示级框架,将每个领域-任务专家建模为共享支撑上隐藏表示的分布。HWM构建任务级和全局Wasserstein重心,以捕获任务内领域变异和跨任务结构,从而通过Wasserstein派生权重实现免训练的专家聚合,或通过混合Wasserstein对齐损失进行基于训练的通才学习。在四个NLP任务(每个任务四个领域)上的实验表明,HWM在MD-MTL设置中实现了优越的有效性和泛化能力。

英文摘要

Multi-domain multi-task learning (MD-MTL) aims to build a single generalist model that performs well across heterogeneous domains and tasks. However, joint training often suffers from interference under distribution shifts. Existing model merging methods mostly operate on model parameters while overlooking the geometric structure of latent representation distributions across domains and tasks. To address these limitations, we propose Hierarchical Wasserstein Merging (HWM), a representation-level framework that models each domain-task specialist as a distribution of hidden representations on a shared support. HWM constructs task-level and global Wasserstein barycenters to capture within-task domain variation and cross-task structure, enabling either training-free specialist aggregation by Wasserstein-derived weights or training-based generalist learning through a hybrid Wasserstein alignment loss. Experiments on four NLP tasks across four domains per task show that HWM achieves superior effectiveness and generalization capability in MD-MTL settings.

发表机构

  • Virginia Tech(弗吉尼亚理工大学)
  • Amazon(亚马逊)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑