Beyond Factual Knowledge: Benchmarking and Learning Step-Level Procedural Rule Reasoning in Large Language Models
超越事实知识:大型语言模型中步骤级过程规则推理的基准测试与学习
Bohan Yu, Pengfei Cao, Chen Han, Chenxi Zhou, Zhiheng Zhang, Zhiyang Xie, Wenhao Teng, Xiangwen Liao, Jun Zhao, Kang Liu
机构
*
University of Chinese Academy of Sciences(中国科学院大学)
;
Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
;
Academy of Mathematics and Systems Science, Chinese Academy of Sciences(中国科学院数学与系统科学研究院)
机构
*
University of California, Irvine(加利福尼亚大学欧文分校)
;
University of California, Los Angeles(加利福尼亚大学洛杉矶分校)
;
Monash University(莫纳什大学)
;
Columbia University(哥伦比亚大学)
;
University of Pennsylvania(宾夕法尼亚大学)
Beyond RGB: Benchmarking and Enhancing MLLMs for Hyperspectral Image Understanding via Training-Free Reasoning Framework
HM-Bench:多模态大语言模型在高光谱遥感中的综合基准
Xinyu Zhang, Zurong Mai, Qingmei Li, Xiaoya Fan, Zjin Liao, Haoyuan Liang, Yibin Wen, Yuhang Chen, Chan Tsz Ho, Bi Tianyuan, Ruifeng Su, Zihao Qiang, Juepeng Zheng, Jianxi Huang, Yutong Lu, Haohuan Fu
机构
*
Sun Yat-sen University(中山大学)
;
Tsinghua Shenzhen International Graduate School(清华大学深圳国际研究生院)
;
China Agricultural University(中国农业大学)
;
Southwest Jiaotong University(西南交通大学)
;
Southwest University(西南大学)
;
National Supercomputing Center in Shenzhen(国家超级计算深圳中心)
Teaching agentic AI to learn expert reasoning for rare disease diagnosis
LiteOdyssey: 一种用于可解释罕见病诊断的轻量级推理AI智能体
Minh-Ha Nguyen, Erica Gray, Bryce A. Schuler, Kevin W. Byram, Chih-Ting Yang, Fan Ma, Hua Xu, Wu-Chen Su, Chao Yan, Wei-Qi Wei, Adam Wright, Lisa Bastarache, Josh F. Peterson, Lingyao Li, Siyuan Ma, Undiagnosed Diseases Network, Rizwan Hamid, Thomas A. Cassini, Cathy Shyr
机构
*
Vanderbilt University(范德堡大学)
;
Vanderbilt University Medical Center(范德堡大学医学中心)
;
University of South Florida(南佛罗里达大学)
The IOL-AI Challenge: An Open Challenge towards Advancing Linguistic Reasoning
IOL-AI挑战赛:一项旨在推进语言推理的开放挑战赛
Eduardo Sánchez, Rita Berrada, Dan-Mircea Mirea, Sara Rajaee, Alexander Piperski, Ana Meta Dolinar, Boris Iomdin, Andrey Nikulin, Mariya Shmatova, Marzieh Fadaee, Julia Kreutzer
机构
*
University College London(伦敦大学学院)
;
Meta
;
CentraleSupélec(中央高等电力学院)
;
McGill University(麦吉尔大学)
;
Cohere Labs(Cohere实验室)
;
Princeton University(普林斯顿大学)
;
University of Amsterdam(阿姆斯特丹大学)
;
Stockholm University(斯德哥尔摩大学)
;
Faculty of Liberal Arts and Sciences(文理学院)
;
Universidade Federal de Goiás(戈亚斯联邦大学)
机构
*
Department of Computer Science and Engineering, Hong Kong University of Science and Technology(香港科技大学计算机科学与工程系)
;
Department of Pathology, Nanfang Hospital, Southern Medical University(南方医科大学南芳医院病理科)
;
Department of Pathology, School of Basic Medical Sciences, Southern Medical University(南方医科大学基础医学学院病理科)
;
Department of Anatomical and Cellular Pathology, Chinese University of Hong Kong(香港中文大学解剖与细胞病理学系)
;
Guangdong Provincial Key Laboratory of Molecular Tumor Pathology(广东省分子肿瘤病理学重点实验室)
;
Jinfeng Laboratory(锦风实验室)
;
Department of Chemical and Biological Engineering, Hong Kong University of Science and Technology(香港科技大学化学与生物工程系)
;
Division of Life Science, Hong Kong University of Science and Technology(香港科技大学生命科学系)
;
State Key Laboratory of Nervous System Disorders, The Hong Kong University of Science and Technology(香港科技大学神经系统疾病国家重点实验室)
;
HKUST Shenzhen-Hong Kong Collaborative Innovation Research Institute, The Hong Kong University of Science and Technology(香港科技大学深圳-香港协同创新研究院)
When Do LLMs Apply the Wrong Law? Diagnosing LLM Failures in Temporal Legal Reasoning
大型语言模型何时适用错误法律?诊断大型语言模型在时间法律推理中的失败
Yiqian Huang, Shuyuan Zheng, Qianying Liu, Shaowen Peng, Yuntao Kong, Kotaro Funakoshi, Chuan Xiao, Manabu Okumura, Yang Cao
机构
*
Institute of Science Tokyo(东京科学大学)
;
Osaka University(大阪大学)
;
NII LLMC(日本信息学研究所语言、语言学与媒体中心)
;
Nara Institute of Science and Technology(奈良科学技术研究所)
;
Center of Juris-Informatics, ROIS-DS(日本学术研究推进机构跨学科科学研究中心法学信息学中心)