Closed-Form Spectral Regularization for Multi-Task Model Merging
多任务模型融合的闭式谱正则化
Yongxian Wei, Runxi Cheng, Xingxuan Zhang, Li Shen, Chun Yuan, Peng Cui, Dacheng Tao
机构
*
Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院)
;
Department of Computer Science and Technology, Tsinghua University(清华大学计算机科学与技术系)
;
Sun Yat-sen University(中山大学)
;
Nanyang Technological University(南洋理工大学)
机构
*
Faculty of Engineering, Shenzhen MSU-BIT University(深圳北理莫斯科大学工程学院)
;
Faculty of CMC, Shenzhen MSU-BIT University(深圳北理莫斯科大学计算机、数学与力学学院)
;
Department of CDS, Indian Institute of Science(印度科学学院计算机与数据科学系)
专题命中
文档图表理解
:multimodal large language model(abstract);分类 cs.CV
A report-grounded vision-language foundation model for colonoscopy from 280000 routine reports
基于28万份常规报告的肠镜报告驱动的视觉-语言基础模型
Jia Yu, Yan Zhu, Yili He, Zilong Wang, Xinyang Jiang, Peiyao Fu, Ruijie Yang, Tianyi Chen, Siyuan Li, Zhihua Wang, Fei Wu, Quanlin Li, Xian Yang, Pinghong Zhou, Shuo Wang
机构
*
Digital Medical Research Center, School of Basic Medical Sciences, Fudan University(复旦大学基础医学院数字医学研究中心)
;
Shanghai Collaborative Innovation Center of Endoscopy(上海内镜诊疗协同创新中心)
;
Zhejiang University(浙江大学)
;
Shanghai Institute for Advanced Study of Zhejiang University(浙江大学上海高等研究院)
;
Alliance Manchester Business School, The University of Manchester(曼彻斯特大学联盟曼彻斯特商学院)
;
Data Science Institute, Imperial College London(伦敦帝国理工学院数据科学研究所)
;
Microsoft Research Asia(微软亚洲研究院)
StructuredEdit: Constraint-Aware Graphic Design Editing via Differentiable Parameter Propagation
StructuredEdit:通过可微参数传播实现约束感知的平面设计编辑
Veeramanohar Avudaiappan, Ritwik Murali
机构
*
Department of Electrical and Electronics Engineering, Amrita School of Engineering, Coimbatore, Amrita Vishwa Vidyapeetham India(电子与电子工程系,阿米特拉工程学院,科伊巴托尔,阿米特拉世界学院,印度)
;
Department of Computer Science and Engineering, Amrita School of Computing, Coimbatore, Amrita Vishwa Vidyapeetham India(计算机科学与工程系,阿米特拉计算学院,科伊巴托尔,阿米特拉世界学院,印度)
CommentsAccepted to the 43rd International Conference on Machine Learning (ICML 2026). 22 pages, 11 figures. Code and dataset available at https://github.com/zimoqingfeng/MORE
机构
*
Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院)
;
SparcAI Inc(SparcAI公司)
;
University of Science and Technology of China(中国科学技术大学)
;
The Hong Kong Polytechnic University(香港理工大学)
;
University of Electronic Science and Technology of China(电子科技大学)
;
Nanyang Technological University(南洋理工大学)
;
Institute for AI Industry Research (AIR), Tsinghua University(清华大学人工智能产业研究院)
Advancing WordArt-Oriented Scene Text Recognition: Datasets and Methods
推进面向艺术字的场景文本识别:数据集与方法
Xingsong Ye, Yongkun Du, Jiaxin Zhang, Haojie Zhang, Chong Sun, Chen Li, Jing Lyu, Zhineng Chen
机构
*
Institute of Trustworthy Embodied AI, Fudan University(复旦大学可信具身人工智能研究所)
;
Shanghai Key Laboratory of Multimodal Embodied AI, Fudan University(复旦大学上海市多模态具身人工智能重点实验室)
;
WeChat Vision, Tencent Inc.(腾讯微信视觉团队)
;
South China University of Technology(华南理工大学)