机构
*
The University of Sydney, Sydney, Australia(悉尼大学)
;
Nanjing University of Science(南京理工大学)
;
Nanjing University, Nanjing, China(南京大学)
;
National University of Singapore, Singapore, Singapore(新加坡国立大学)
From Scores to Steps: Diagnosing and Improving LLM Performance in Evidence-Based Medical Calculations
从分数到步骤:诊断和改进证据医学计算中LLM的性能
Benlu Wang, Iris Xia, Yifan Zhang, Junda Wang, Feiyun Ouyang, Shuo Han, Arman Cohan, Hong Yu, Zonghai Yao
机构
*
Department of Computer Science, Yale University, CT, USA(耶鲁大学计算机科学系)
;
Center for Healthcare Organization and Implementation Research, VA Bedford Health Care(VA贝福德医疗中心健康组织与实施研究中心)
;
Miner School of Computer and Information Sciences, UMass Lowell, MA, USA(UMass洛厄尔矿尔计算机与信息科学学院)
;
Manning College of Information and Computer Sciences, UMass Amherst, MA, USA(UMass阿默斯特马宁信息与计算机科学学院)
CommentsEqual contribution for the first two authors. To appear as an Oral presentation in the proceedings of the Main Conference on Empirical Methods in Natural Language Processing (EMNLP) 2025
EvalQReason: A Framework for Step-Level Reasoning Evaluation in Large Language Models
EvalQReason: 一种用于大型语言模型中分步推理评估的框架
Shaima Ahmad Freja, Ferhat Ozgur Catak, Betul Yurdem, Chunming Rong
机构
*
Department of Electrical Engineering and Computer Science, University of Stavanger(电子工程与计算机科学系,斯塔万格大学)
;
Department of Electrical and Electronics Engineering, Izmir Bakircay University(电子与电气工程系,伊兹密尔巴克伊大学)
Benchmarking neural surrogates on realistic spatiotemporal multiphysics flows
在现实时空多物理场流中评估神经替代模型
Runze Mao, Rui Zhang, Xuan Bai, Tianhao Wu, Teng Zhang, Zhenyi Chen, Minqi Lin, Bocheng Zeng, Yangchen Xu, Yingxuan Xiang, Haoze Zhang, Shubham Goswami, Pierre A. Dawe, Yifan Xu, Zhenhua An, Mengtao Yan, Xiaoyi Lu, Yi Wang, Rongbo Bai, Haobu Gao, Xiaohang Fang, Han Li, Hao Sun, Zhi X. Chen
机构
*
State Key Laboratory for Turbulence and Complex Systems, School of Mechanics and Engineering Science, Peking University(湍流与复杂系统国家重点实验室,力学与工程科学学院,北京大学)
;
Gaoling School of Artificial Intelligence, Renmin University of China(Gallagher人工智能学院,中国人民大学)
;
AI for Science Institute(人工智能科学研究院)
;
University of Calgary(卡尔加里大学)
;
Kyoto University(京都大学)
;
FM Global(FM全球)
;
LandSpace Technology Corporation Ltd.(陆地方向技术有限公司)
;
Aero Engine Academy of China(中国航空发动机学院)
T-LLM: Teaching Large Language Models to Forecast Time Series via Temporal Distillation
T-LLM:通过时间蒸馏教大语言模型进行时间序列预测
Suhan Guo, Bingxu Wang, Shaodan Zhang, Furao Shen
机构
*
State Key Laboratory for Novel Software Technology, Nanjing University(南京大学新型软件技术国家重点实验室)
;
School of Artificial Intelligence, Nanjing University(南京大学人工智能学院)
AgroFlux: A Spatial-Temporal Benchmark for Carbon and Nitrogen Flux Prediction in Agricultural Ecosystems
AgroFlux:农业生态系统碳和氮通量预测的空间-时间基准
Qi Cheng, Licheng Liu, Yao Zhang, Mu Hong, Yiqun Xie, Xiaowei Jia
机构
*
University of Pittsburgh(匹兹堡大学)
;
University of Minnesota - Twin Cities(明尼苏达大学-双城分校)
;
Colorado State University(科罗拉多州立大学)
;
University of Maryland(马里兰大学)
Eliciting Trustworthiness Priors of Large Language Models via Economic Games
通过经济游戏 eliciting 大语言模型的信任度先验
Siyu Yan, Lusha Zhu, Jian-Qiao Zhu
机构
*
University of Hong Kong(香港大学)
;
The University of Hong Kong(香港大学)
;
Peking University(北京大学)
;
School of Psychological and Cognitive Sciences, Peking University(北京大学心理与认知科学学院)
;
Beijing Key Laboratory of Behavior and Mental Health, Peking University(北京大学行为与心理健康重点实验室)
;
IDG/McGovern Institute for Brain Research, Peking University(北京大学脑科学研究院)
;
Peking-Tsinghua Center for Life Sciences, Peking University(北京大学-清华大学生命科学中心)
;
Key Laboratory of Machine Perception, Ministry of Education, China(教育部机器感知重点实验室)
机构
*
LARG, Research Center for Social Computing and Interactive Robotics, HIT(LARG,社会计算与交互机器人研究中心,哈尔滨工业大学)
;
School of Computer Science and Engineering, Central South University(计算机科学与工程学院,中南大学)
;
iFLYTEK
CommentsFirst Hindi sycophancy benchmark using a three-condition design separating language and cultural effects, with empirical evaluation across four instruction-tuned models
A Survey of AI Methods for Geometry Preparation and Mesh Generation in Engineering Simulation
工程仿真中几何准备和网格生成的AI方法综述
Steven Owen, Nathan Brown, Nikos Chrisochoides, Rao Garimella, Xianfeng Gu, Franck Ledoux, Na Lei, Roshan Quadros, Navamita Ray, Nicolas Winovich, Yongjie Jessica Zhang
机构
*
Sandia National Laboratories(桑迪亚国家实验室)
;
Old Dominion University(旧 Dominion 大学)
;
Los Alamos National Laboratory(洛斯阿拉莫斯国家实验室)
;
New York University / Stony Brook University(纽约大学 / 斯通布鲁克大学)
;
CEA(法国原子能委员会)
;
Dalian University of Technology(大连理工大学)
;
Carnegie Mellon University(卡内基梅隆大学)
机构
*
Zhejiang University, China(浙江大学)
;
Westlake University, China(西湖大学)
;
Beijing Innovation Center of Humanoid Robotics, China(北京人形机器人创新中心)
;
University of Manchester, UK(曼彻斯特大学)
;
Southern University of Science(南方科技大学)
;
Xiamen University Malaysia, Malaysia(厦门大学马来西亚分校)
机构
*
Civil and Environmental Engineering, University of Maryland(大学环境工程系,马里兰州大学)
;
Computer and Information Science, University of Pennsylvania(计算机与信息科学系,宾夕法尼亚大学)
;
Statistics, University of California, Berkeley(统计学系,加州大学伯克利分校)
Model Specific Task Similarity for Vision Language Model Selection via Layer Conductance
为视觉语言模型选择而基于模型特定任务相似性的层导电性
Wei Yang, Hong Xie, Tao Tan, Xin Li, Defu Lian, Enhong Chen
机构
*
School of Computer Science and Technology, University of Science and Technology of China(计算机科学与技术学院,科学技术大学)
;
School of AI and Data Science, University of Science and Technology of China(人工智能与数据科学学院,科学技术大学)