Scaling Laws in the Tiny Regime: How Small Models Change Their Mistakes
小规模 regime 中的缩放规律:小型模型如何改变其错误
Mohammed Alnemari, Rizwan Qureshi, Nader Begrazadah
机构
*
Faculty of Computer Science and Information Technology, University of Tabuk(计算机科学与信息科技学院,塔布克大学)
;
AIST Research Center, University of Tabuk(AIST研究中心,塔布克大学)
;
Department of Computer Science, Salim Habib University(计算机科学系,Salim Habib大学)
;
University of California, Irvine(加州大学尔湾分校)
专题命中
预训练与数据
:large language model(abstract);language model(abstract);分类 cs.AI、cs.LG
SD-MoE: Spectral Decomposition for Effective Expert Specialization
SD-MoE:通过谱分解实现有效的专家专业化
Ruijun Huang, Fang Dong, Xin Zhang, Hengjie Cao, Zhendong Huang, Anrui Chen, Jixian Zhou, Mengyi Chen, Yifeng Yang, Mingzhi Dong, Yujiang Wang, Jinlong Hou, Qin Lv, Robert P. Dick, Yuan Cheng, Fan Yang, Tun Lu, Chun Zhang, Li Shang
机构
*
College of Computer Science and Artificial Intelligence, Fudan University, Shanghai, China(复旦大学计算机科学与人工智能学院)
;
University of Bath, Bath, United Kingdom(巴斯大学)
;
Oxford Suzhou Centre for Advanced Research, Suzhou, China(牛津苏泽研究中心)
;
Department of Electrical Engineering and Computer Science, University of Michigan(密歇根大学电气工程与计算机科学系)
;
Shanghai Innovation Institute, Shanghai, China(上海创新研究院)
;
Department of Computer Science, University of Colorado Boulder, Colorado, USA(科罗拉多大学博尔德分校计算机科学系)
;
Research Institute of Tsinghua University in Shenzhen, Shenzhen, China(清华大学深圳研究院)
;
Greater Bay Area National Center of Technology Innovation, Research Institute of Tsinghua University in Shenzhen, Shenzhen, China(粤港澳大湾区国家技术创新中心,清华大学深圳研究院)
;
School of Microelectronics, Fudan University, Shanghai, China(复旦大学微电子学院)
专题命中
预训练与数据
:large language model(abstract);language model(abstract);分类 cs.AI、cs.LG
EVA: Towards a universal model of the immune system
EVA:朝着免疫系统通用模型迈进
Scienta Team, Ethan Bandasack, Vincent Bouget, Apolline Bruley, Yannis Cattan, Charlotte Claye, Matthew Corney, Julien Duquesne, Karim El Kanbi, Aziz Fouché, Pierre Marschall, Francesco Strozzi
Data Kernel Perspective Space Performance Guarantees for Synthetic Data from Transformer Models
变换器模型合成数据的Data Kernel视角空间性能保证
Michael Browder, Kevin Duh, J. David Harris, Vince Lyzinski, Paul McNamee, Youngser Park, Carey E. Priebe, Peter Viechnicki
机构
*
Department of Mathematics at the University of Maryland, College Park(马里兰大学College Park数学系)
;
Human Language Technology Center of Excellence, Johns Hopkins University(约翰霍普金斯大学人机语言技术中心)
;
Center for Imaging Science (CIS), the Institute for Computational Medicine (ICM), and the Mathematical Institute for Data Science (MINDS), Johns Hopkins University(约翰霍普金斯大学影像科学中心(CIS)、计算医学研究所(ICM)和数据科学数学研究所(MINDS))
;
Department of Applied Mathematics and Statistics (AMS), the Center for Imaging Science (CIS), and the Mathematical Institute for Data Science (MINDS), Johns Hopkins University(约翰霍普金斯大学应用数学与统计学系(AMS)、影像科学中心(CIS)和数据科学数学研究所(MINDS))
Do Transformers Have the Ability for Periodicity Generalization?
Transformer 是否具备周期性泛化能力?
Huanyu Liu, Ge Li, Yihong Dong, Sihan Wu, Peixu Wang, Sihao Cheng, Taozhi Chen, Kechi Zhang, Hao Zhu, Tongxuan Liu
机构
*
School of Computer Science, Peking University(北京大学计算机科学学院)
;
College of AI, Tsinghua University(清华大学人工智能学院)
;
School of Science and Engineering, The Chinese University of Hong Kong (Shenzhen)(香港中文大学(深圳)科学与工程学院)
专题命中
预训练与数据
:large language model(abstract);language model(abstract);分类 cs.AI、cs.LG