arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校)

2026-02-03 至 2026-02-03 共收录 4
2602.01410 2026-02-03 cs.LG cs.AR

SNIP: An Adaptive Mixed Precision Framework for Subbyte Large Language Model Training

SNIP:一种用于子字节大语言模型训练的自适应混合精度框架

Yunjie Pan, Yongyi Yang, Hanmei Yang, Scott Mahlke

机构 * University of Michigan(密歇根大学) NTT Research, Inc.(NTT研究公司) University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校)

AI总结 SNIP通过自适应混合精度框架,有效提升大语言模型训练效率,减少FLOPs达80%并保持模型质量。

Comments Accepted to ASPLOS 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.01239 2026-02-03 cs.CL cs.IR

Inferential Question Answering

推断性问答

Jamshid Mozafari, Hamed Zamani, Guido Zuccon, Adam Jatowt

机构 * University of Innsbruck(因斯布鲁克大学) University of Massachusetts Amherst(马萨诸塞大学阿姆赫斯特分校) The University of Queensland(昆士兰大学)

AI总结 本文提出推断性QA任务,通过构建QUIT数据集,发现传统QA方法在推断任务中表现不佳,揭示当前QA流程难以处理基于推断的推理。

Comments Proceedings of the ACM Web Conference 2026 (WWW 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00970 2026-02-03 cs.CL cs.GT

Verification Required: The Impact of Information Credibility on AI Persuasion

验证所需:信息可信度对AI说服力的影响

Saaduddin Mahmud, Eugene Bagdasarian, Shlomo Zilberstein

机构 * Manning College of Information and Computer Sciences, University of Massachusetts Amherst, Massachusetts, USA(信息与计算机科学学院,马萨诸塞大学阿默斯特分校,马萨诸塞州,美国)

AI总结 研究提出MixTalk游戏模型,通过验证与不可验证声明的结合,探讨信息可信度对AI说服力的影响,并提出TOPD方法提升接收者鲁棒性。

Comments 19 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16584 2026-02-03 cs.CL cs.AI

From Scores to Steps: Diagnosing and Improving LLM Performance in Evidence-Based Medical Calculations

从分数到步骤:诊断和改进证据医学计算中LLM的性能

Benlu Wang, Iris Xia, Yifan Zhang, Junda Wang, Feiyun Ouyang, Shuo Han, Arman Cohan, Hong Yu, Zonghai Yao

机构 * Department of Computer Science, Yale University, CT, USA(耶鲁大学计算机科学系) Center for Healthcare Organization and Implementation Research, VA Bedford Health Care(VA贝福德医疗中心健康组织与实施研究中心) Miner School of Computer and Information Sciences, UMass Lowell, MA, USA(UMass洛厄尔矿尔计算机与信息科学学院) Manning College of Information and Computer Sciences, UMass Amherst, MA, USA(UMass阿默斯特马宁信息与计算机科学学院)

AI总结 本文提出MedRaC框架,通过分步评估和代码执行提升LLM在证据医学计算中的准确性,揭示现有评估方法的不足,并推动临床可信度的提升。

Comments Equal contribution for the first two authors. To appear as an Oral presentation in the proceedings of the Main Conference on Empirical Methods in Natural Language Processing (EMNLP) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏