Commentsv5: Fixed PDF rendering compatibility issue affecting Apple PDFKit (macOS Preview/iOS PDF viewer). No changes to technical content compared to v4
机构
*
Beijing University of Posts and Telecommunications(北京邮电大学)
;
Tsinghua University(清华大学)
;
Institute of Artificial Intelligence (TeleAI)(人工智能研究所)
;
ShiFang Technology Inc.(ShiFang科技公司)
Wasserstein Distributionally Robust Regret Optimization for Reinforcement Learning from Human Feedback
Wasserstein分布鲁棒遗憾优化用于人类反馈的强化学习
Yikai Wang, Shang Liu, Jose Blanchet
机构
*
Department of Statistics and Operations Research, University of North Carolina(统计与运筹学系,北卡罗来纳大学)
;
Imperial Business School, Imperial College London(帝国理工学院伦敦商学院)
;
Department of Management Science and Engineering, Stanford University(管理科学与工程系,斯坦福大学)
机构
*
Vellore Institute of Technology(维洛雷理工学院)
;
University of Massachusetts Amherst(马萨诸塞大学阿姆赫斯特分校)
;
Northwestern University(西北大学)
;
Yale University(耶鲁大学)
;
Algoverse AI Research(Algoverse AI研究)
Comments12 pages, 4 figures, 6 tables. Includes ablation study across Qwen2.5-7B-Instruct and Llama-3.1-8B-Instruct on 5 math reasoning benchmarks (GSM8K, MATH500, Minerva, AIME24, Gaokao2023). GPT-4.1 used for structured evaluation of reasoning quality
机构
*
University of Arizona, USA(亚利桑那大学)
;
Arizona State University, USA(亚利桑那州立大学)
;
Now at Google LLC, work done at Rice University(现就职于谷歌公司,曾就职于里士大学)
;
Clemson University, USA(克莱姆森大学)
;
Washington University in St. Louis, USA(圣路易斯华盛顿大学)
;
Halmstad University, Sweden(哈姆斯塔德大学)
;
Guangdong Institute of Intelligence Science and Technology, China(广东智能科学与技术研究院)
Self-Prompting Small Language Models for Privacy-Sensitive Clinical Information Extraction
面向隐私敏感的临床信息抽取的自提示小型语言模型
Yao-Shun Chuang, Tushti Mody, Uday Pratap Singh, Shirindokht Shiraz, Chun-Teh Lee, Ryan Brandon, Muhammad F Walji, Xiaoqian Jiang, Bunmi Tokede
机构
*
McWilliams School of Biomedical Informatics, The University of Texas Health Science Center at Houston(德克萨斯大学健康科学中心休斯顿分校麦克威廉斯生物医学信息学学院)
;
School of Public Health, The University of Texas Health Science Center at Houston(德克萨斯大学健康科学中心休斯顿分校公共卫生学院)
;
School of Dentistry, The University of Texas Health Science Center at Houston(德克萨斯大学健康科学中心休斯顿分校牙科学院)
;
Willamette Dental and Skourtes Institute(威廉特牙科与斯库尔特斯研究所)
Truthful or Fabricated? Using Causal Attribution to Mitigate Reward Hacking in Explanations
真实还是编造?利用因果归因减轻解释中的奖励作弊
Pedro Ferreira, Wilker Aziz, Ivan Titov
机构
*
Institute for Logic, Language and Computation (ILLC), University of Amsterdam(逻辑、语言与计算研究所(ILLC),阿姆斯特丹大学)
;
Institute for Language, Cognition and Computation (ILCC), University of Edinburgh(语言、认知与计算研究所(ILCC),爱丁堡大学)
LearNAT: Learning NL2SQL with AST-guided Task Decomposition for Large Language Models
LearNAT: 基于AST引导任务分解的NL2SQL大语言模型学习
Weibin Liao, Xin Gao, Tianyu Jia, Rihong Qiu, Yifan Zhu, Yang Lin, Xinyu Ma, Junfeng Zhao, Yasha Wang
机构
*
School of Computer Science, Peking University(北京大学计算机科学系)
;
Key Laboratory of High Confidence Software Technologies, Ministry of Education(教育部高可信软件技术重点实验室)
;
Big Data Technology Research Center, Nanhu Laboratory(纳米实验室大数据技术研究中心)
;
National Engineering Research Center For Software Engineering, Peking University(北京大学软件工程国家工程研究中心)
;
Peking University Information Technology Institute (Tianjin Binhai)(北京大学信息技术研究院(天津滨海))
;
School of Computer Sciences, Beijing University of Posts and Telecommunications(北京邮电大学计算机科学系)
;
Huawei Technologies Co., Ltd(华为技术有限公司)
;
Seed, ByteDance Inc.(字节跳动公司)