arXivDaily arXiv每日学术速递 周一至周五更新

期刊&会议

International Conference on Machine Learning · 会议 · Machine Learning

2025-12-02 至 2025-12-02 共收录 4
2410.11842 2025-12-02 cs.CV cs.AI cs.LG

MoH: Multi-Head Attention as Mixture-of-Head Attention

MoH:多头注意力作为专家混合注意力

Peng Jin, Bo Zhu, Li Yuan, Shuicheng Yan

机构 * School of Electronic and Computer Engineering, Shenzhen Graduate School, Peking University, Shenzhen, China(电子与计算机工程系,深圳研究生院,北京大学,深圳,中国) Pengcheng Laboratory, Shenzhen, China(鹏城实验室,深圳,中国) School of AI for Science, Shenzhen Graduate School, Peking University, Shenzhen, China(科学人工智能学院,深圳研究生院,北京大学,深圳,中国) National University of Singapore, Singapore(新加坡国立大学,新加坡)

AI总结 MoH通过将注意力头视为专家,提升推理效率并优化性能,仅使用部分注意力头即可超越传统多头注意力。

Comments Accepted by ICML 2025, code: https://github.com/SkyworkAI/MoH

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10918 2025-12-02 cs.LG cs.AI

ToMA: Token Merge with Attention for Diffusion Models

ToMA: 通过注意力机制的令牌合并用于扩散模型

Wenbo Lu, Shaoyi Zheng, Yuxuan Xia, Shengjie Wang

机构 * Department of Computer Science, New York University(纽约大学计算机科学系)

AI总结 ToMA通过重新设计令牌合并方法,提升扩散模型的GPU效率,减少生成延迟并优化实际性能。

Comments In proceedings of the 42nd International Conference on Machine Learning (ICML 2025). Code available at https://github.com/wenboluu/ToMA

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.10205 2025-12-02 cs.LG

AWP: Activation-Aware Weight Pruning and Quantization with Projected Gradient Descent

AWP: 基于激活感知的权重剪枝与量化方法(投影梯度下降)

Jing Liu, Toshiaki Koike-Akino, Ye Wang, Hassan Mansour, Matthew Brand

机构 * Mitsubishi Electric Research Laboratories (MERL)(三菱电机研究实验室)

AI总结 AWP通过投影梯度下降方法实现激活感知的权重剪枝与量化,优于现有LLM压缩技术。

Comments ICML 2025 workshop on Efficient Systems for Foundation Models

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19097 2025-12-02 cs.LG stat.ML

Towards Robust Influence Functions with Flat Validation Minima

迈向具有平坦验证极小值的鲁棒影响函数

Xichen Ye, Yifan Wu, Weizhong Zhang, Cheng Jin, Yifan Chen

机构 * Fudan University(复旦大学) Hong Kong Baptist University(香港 Baptist 大学) Shanghai Key Laboratory of Intelligent Information Processing(上海智能信息处理关键实验室) Innovation Center of Calligraphy and Painting Creation Technology(书法和绘画创作技术创新中心)

AI总结 本文提出了一种针对平坦验证极小值的影响函数估计方法,解决了深度学习中因验证风险尖锐性导致的影响估计不准确问题。

Comments Accepted by ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏