arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

Johns Hopkins University(约翰斯·霍普金斯大学)

2025-12-18 至 2025-12-18 共收录 3
2505.17083 2025-12-18 cs.CL cs.LG stat.ML

Scale-invariant Attention

尺度不变注意力

Ben Anson, Xi Wang, Laurence Aitchison

机构 * School of Mathematics University of Bristol(布里斯托大学数学学院) Department of Computer Science Johns Hopkins University(约翰霍普金斯大学计算机科学系) School of Computer Science University of Bristol(布里斯托大学计算机科学学院)

AI总结 本文提出了一种尺度不变的注意力机制,通过简单的位置依赖变换实现总注意力和稀疏性的尺度不变性,提升了长上下文推理和检索性能。

Comments Accepted at Neurips 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.03625 2025-12-18 math.OC cs.LG

Deterministic Global Optimization of the Acquisition Function in Bayesian Optimization: To Do or Not To Do?

贝叶斯优化中获取函数确定性全局优化的探讨:做或不做?

Anastasia Georgiou, Daniel Jungen, Luise Kaven, Verena Hunstig, Constantine Frangakis, Ioannis Kevrekidis, Alexander Mitsos

机构 * Chemical & Biomolecular Engineering, Johns Hopkins University, Baltimore, MD 21218, USA(约翰霍普金斯大学化学与生物分子工程系) Process Systems Engineering (AVT.SVT), RWTH Aachen University, Aachen, Germany(亚琛工业大学过程系统工程系) Department of Biostatistics, Bloomberg School of Public Health, Johns Hopkins University, Baltimore, MD 21218, USA(约翰霍普金斯大学比尔·盖茨公共卫生学院生物统计学系) Department of Medicine, Johns Hopkins University, Baltimore, MD 21218, USA(约翰霍普金斯大学医学系) Applied Mathematics & Statistics, Johns Hopkins University, Baltimore, MD 21218, USA(约翰霍普金斯大学应用数学与统计学系) JARA-CSD, 52056 Aachen, Germany(亚琛大学JARA-CSD研究中心) Institute of Climate and Energy Systems, Energy Systems Engineering (ICE-1), Forschungszentrum Jülich GmbH, 52425 Jülich, Germany(吕贝克研究中心气候与能源系统研究所)

AI总结 本文探讨了在贝叶斯优化中使用确定性全局求解器MAiNGO优化获取函数的优劣,发现其在特定条件下可能更优或更劣,取决于获取函数的探索与利用倾向。

Comments 39 pages, 8 figures, 11 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18148 2025-12-18 cs.CL cs.AI cs.LG

Hidden in the Haystack: Smaller Needles are More Difficult for LLMs to Find

haystack中隐藏的针:更小的针对LLM来说更难找到

Owen Bianchi, Mathew J. Koretsky, Maya Willey, Chelsea X. Alvarado, Tanay Nayak, Adi Asija, Nicole Kuznetsov, Mike A. Nalls, Faraz Faghri, Daniel Khashabi

机构 * Center for Alzheimer’s Disease and Related Dementias, NIA, NIH(阿尔茨海默病及相关痴呆症研究中心,国家老龄化研究所,国家卫生研究院) DataTecnica LLC(DataTecnica公司) Johns Hopkins University(约翰霍普金斯大学) Laboratory of Neurogenetics, NIA, NIH(神经遗传学实验室,国家老龄化研究所,国家卫生研究院)

AI总结 本文研究了黄金上下文大小对LLM长上下文问答性能的影响,发现较短的黄金上下文会显著降低模型性能,揭示了上下文长度对模型表现的关键作用。

Comments Under Review

详情

展开后加载摘要…

URL PDF HTML 收藏