MorphoBench: A Benchmark with Difficulty Adaptive to Model Reasoning
Xukai Wang, Xuanbo Liu, Mingrui Chen, Haitian Zhong, Xuanlin Yang, Bohan Zeng, Jinbo Hu, Hao Liang, Junbo Niu, Xuchen Li, Ruitao Wu, Ruichuan An, Yang Shi, Liu Liu, Xu-Yao Zhang, Qiang Liu, Zhouchen Lin, Wentao Zhang, Bin Dong
机构
*
Zhongguancun Academy(中关村学院)
;
Peking University(北京大学)
;
Beihang University(北航)
;
School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)
A Clinically-Grounded Two-Stage Framework for Renal CT Report Generation
Renjie Liang, Zhengkang Fan, Jinqian Pan, Chenkun Sun, Bruce Daniel Steinberg, Russell Terry, Jie Xu
机构
*
Department of Health Outcomes and Biomedical Informatics, University of Florida(健康结果与生物医学信息学系,佛罗里达大学)
;
Department of Urology, University of Florida(泌尿外科系,佛罗里达大学)
MetaBench: A Multi-task Benchmark for Assessing LLMs in Metabolomics
Yuxing Lu, Xukai Zhao, J. Ben Tamo, Micky C. Nnamdi, Rui Peng, Shuang Zeng, Xingyu Hu, Jinzhuo Wang, May D. Wang
机构
*
Wallace H. Coulter Department of Biomedical Engineering, Georgia Institute of Technology and Emory University(沃克生物医学工程部门,佐治亚理工学院和埃默里大学)
;
College of Future of Technology, Peking University(未来技术学院,北京大学)
;
School of Architecture, Tsinghua University(建筑学院,清华大学)
;
School of Electrical and Computer Engineering, Georgia Institute of Technology(电气与计算机工程学院,佐治亚理工学院)
;
School of Computer Science, Georgia Institute of Technology(计算机科学学院,佐治亚理工学院)
Edoardo Loru, Jacopo Nudo, Niccolò Di Marco, Alessandro Santirocchi, Roberto Atzeni, Matteo Cinelli, Vincenzo Cestari, Clelia Rossi-Arnaud, Walter Quattrociocchi
机构
*
Department of Computer, Control and Management Engineering, Sapienza University of Rome(计算机、控制与管理工程系,罗马萨皮恩扎大学)
;
Department of Computer Science, Sapienza University of Rome(计算机科学系,罗马萨皮恩扎大学)
;
Department of Legal, Social, and Educational Sciences, Tuscia University(法律、社会与教育科学系,图斯西亚大学)
;
Department of Psychology, Sapienza University of Rome(心理学系,罗马萨皮恩扎大学)
Robust or Suggestible? Exploring Non-Clinical Induction in LLM Drug-Safety Decisions
Siying Liu, Shisheng Zhang, Indu Bala
专题命中
推理评测
:reasoning(abstract);分类 cs.CL
CommentsPreprint of a paper accepted as a poster at the NeurIPS 2025 Workshop on Generative AI for Health (GenAI4Health). The final camera-ready workshop version may differ. Licensed under CC BY 4.0
Leveraging LLMs, IDEs, and Semantic Embeddings for Automated Move Method Refactoring
Abhiram Bellur, Fraol Batole, Mohammed Raihan Ullah, Malinda Dilhara, Yaroslav Zharov, Timofey Bryksin, Kai Ishikawa, Haifeng Chen, Masaharu Morimoto, Shota Motoura, Takeo Hosomi, Tien N. Nguyen, Hridesh Rajan, Nikolaos Tsantalis, Danny Dig
机构
*
University of Colorado(科罗拉多大学)
;
Tulane University(路易斯安那州立大学)
;
Amazon Web Services(亚马逊网络服务)
;
JetBrains Research(JetBrains研究)
;
NEC Corporation(日本电报电话公司)
;
NEC Laboratories America(日本电报电话美洲实验室)
;
University of Texas at Dallas(德克萨斯大学达拉斯分校)
;
Concordia University(康科迪亚大学)
;
University of Colorado, JetBrains Research(科罗拉多大学,JetBrains研究)
专题命中
推理评测
:reasoning(abstract);分类 cs.AI
CommentsPublished at the International Conference on Software Maintenance and Evolution (ICSME'25)
Eyes Wide Open: Ego Proactive Video-LLM for Streaming Video
Yulin Zhang, Cheng Shi, Yang Wang, Sibei Yang
机构
*
ShanghaiTech University(上海科技大学)
;
School of Computer Science and Engineering, Sun Yat-sen University(中山大学计算机科学与工程学院)
;
School of Computing and Data Science, The University of Hong Kong(香港大学计算科学与数据科学学院)
专题命中
推理评测
:reasoning(abstract)
CommentsAccepted at NeurIPS 2025 (preview; camera-ready in preparation)
机构
*
Beijing Innovation Center of Humanoid Robotics(北京人形机器人创新中心)
;
State Key Laboratory of Multimedia Information Processing(国家多媒体信息处理重点实验室)
;
School of Computer Science, Peking University(北京大学计算机学院)
;
Hong Kong University of Science and Technology(香港科技大学)