arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

New York University(纽约大学)

2026-08-13 至 2026-08-13 共收录 4
2608.11859 2026-08-13 cs.LG 新提交

Small-Scale Experiments: Are We There Yet?

小规模实验:我们达到目标了吗?

Nicholas Lourie, Kyunghyun Cho, Karen Ullrich, Sanae Lotfi

机构 * FAIR at MSL Meta New York University(纽约大学)

AI总结 该研究指出小规模模型对超参数的敏感性是缩放定律在小规模实验中失效的关键,开发了以模型为中心的研究方法论,通过小规模实验证实Transformer中预归一化随模型规模增大效果更好,表明小规模实验可兑现缩放定律的承诺。

Comments 29 pages, 17 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.11746 2026-08-13 cs.LG cs.CL 新提交

Epiplexity Guided Data Selection and Generation for Out-of-Distribution Generalization

面向分布外泛化的、受Epiplexity引导的数据选择与生成

Ellen Su, Andres Potapczynski, Shikai Qiu, Edward Hughes, Andrew Gordon Wilson

机构 * New York University(纽约大学)

AI总结 该研究提出以Epiplexity为引导,通过数据选择与合成数据生成提升模型分布外泛化能力,实验显示更高Epiplexity可改善零样本及微调任务的下游性能。

Comments Code available for EpiSelect (https://github.com/eysu35/EpiSelect) and EpiGen (https://github.com/eysu35/EpiGen)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.11612 2026-08-13 cs.LG cs.AI 新提交

Dion3: Full-Stack Orthogonal Updates

Dion3:全栈正交更新

Noah Amsel, Jack Zhang, Kwangjun Ahn, Ali Naeimi, Austin Feng, Berlin Chen, Tri Dao, John Langford

机构 * Microsoft Research(微软研究院) New York University(纽约大学) Princeton University(普林斯顿大学) NVIDIA(英伟达) Yale University(耶鲁大学)

AI总结 Dion3是针对Muon优化器开销的改进版本,通过全栈优化降低了正交化、通信等开销,步长时间最多降6倍,可作为Muon的即插即用替代方案。

Comments 37 pages, 23 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.28854 2026-08-13 cs.CL cs.LG q-bio.NC 版本更新

Large language models reorganize representational geometry during in-context learning

大型语言模型在上下文学习中重组表征几何结构

Hua-Dong Xiong, Li Ji-An, Robert C. Wilson, Kwonjoon Lee, Xue-Xin Wei

机构 * School of Psychological and Brain Sciences, Georgia Tech(佐治亚理工学院心理与脑科学学院) Department of Psychology, New York University(纽约大学心理学系) Center of Excellence for Computational Cognition, Georgia Tech(佐治亚理工学院计算认知卓越中心) Honda Research Institute(本田研究院) Departments of Neuroscience and Psychology, The University of Texas at Austin(德克萨斯大学奥斯汀分校神经科学与心理学系)

AI总结 研究大型语言模型在上下文学习中的表征几何重组,发现其性能与任务表征结构相关,并通过原型算法动态调整表征以提高可分性。

Comments Published as a conference paper at COLM 2026

详情

展开后加载摘要…

URL PDF HTML 收藏