MEBench: Benchmarking Large Language Models for Cross-Document Multi-Entity Question Answering
Teng Lin, Yuyu Luo, Honglin Zhang, Jicheng Zhang, Chunlin Liu, Kaishun Wu, Nan Tang
机构
*
The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州))
;
The Hong Kong University of Science and Technology(香港科学与技术大学)
;
China Mobile Information Technology Company Limited(中国移动信息技术有限公司)
专题命中
评测与基准
:large language model(title,abstract);language model(title,abstract);LLM(abstract);分类 cs.CL
机构
*
Chinese University of Hong Kong(中国香港大学)
;
City University of Hong Kong(香港城市大学)
;
University of Oxford(牛津大学)
;
Shanghai AI Laboratory(上海人工智能实验室)
;
The Sixth Affiliated Hospital, Sun Yat-sen University(中山大学第六附属医院)
专题命中
评测与基准
:large language model(title,abstract);language model(title,abstract);prompting(abstract)
Comments40 pages, 22 figures; Accepted by NeurIPS 2025 Dataset and Benchmark Track
机构
*
Shanghai Key Laboratory of Data Science, School of Computer Science, Fudan University(上海数据科学 key laboratory,计算机科学学院,复旦大学)
;
School of Management, Fudan University(管理学院,复旦大学)
;
Tencent Weixin Group(腾讯微信集团)
;
IFM Lab, University of California, Davis(加州大学戴维斯分校IFM实验室)
专题命中
评测与基准
:large language model(title,abstract);language model(title,abstract);分类 cs.CL、cs.AI
Large Language Models for Multilingual Previously Fact-Checked Claim Detection
Ivan Vykopal, Matúš Pikuliak, Simon Ostermann, Tatiana Anikina, Michal Gregor, Marián Šimko
机构
*
Faculty of Information Technology, Brno University of Technology(布拉格技术大学信息学院)
;
Kempelen Institute of Intelligent Technologies(智能技术研究所)
;
German Research Center for Artificial Intelligence (DFKI)(德国人工智能研究中心)
;
Centre for European Research in Trusted AI (CERTAIN)(可信人工智能欧洲研究中心)
专题命中
评测与基准
:large language model(title,abstract);language model(title,abstract);分类 cs.CL
How Model Size, Temperature, and Prompt Style Affect LLM-Human Assessment Score Alignment
Julie Jung, Max Lu, Sina Chole Benker, Dogus Darici
机构
*
Harvard Graduate School of Education(哈佛教育研究生院)
;
Munster University(穆恩斯特大学)
;
Institute of Anatomy and Neurobiology, University of Münster(解剖与神经生物学研究所,穆恩斯特大学)
专题命中
评测与基准
:LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL
Comments9 pages, 4 figures, accepted at NCME AIME 2025
DSA, AIA, and LLMs: Approaches to conceptualizing and auditing moderation in LLM-based chatbots across languages and interfaces in the electoral contexts
Comments16 pages, 7 figures. Accepted for presentation at the 39th Conference on Neural Information Processing Systems (NeurIPS 2025) Workshop on the Foundations of Reasoning in Language Models (FoRLM)
机构
*
Department of Computer Science and Information Engineering(计算机科学与信息工程系)
;
National Taiwan University(国立台湾大学)
;
Institute of Information Science, Academia Sinica(学术院信息研究所)
;
AI Research Center (AINTU)(人工智能研究中心)
专题命中
评测与基准
:LLM(title,abstract);分类 cs.CL、cs.AI
CommentsAccepted as a long findings paper at EMNLP 2025
A Foundation Chemical Language Model for Comprehensive Fragment-Based Drug Discovery
Alexander Ho, Sukyeong Lee, Francis T. F. Tsai
机构
*
Advanced Technology Core for Macromolecular X-Ray Crystallography(宏分子X射线衍射先进技术核心)
;
Verna and Marrs McLean Dept. of Biochemistry and Molecular Pharmacology(韦纳和玛里斯麦克莱恩生物化学与分子药理学系)
;
Dept. of Molecular & Cellular Biology(分子与细胞生物学系)
;
Dept. of Molecular Virology & Microbiology(分子病毒学与微生物学系)
;
Baylor College of Medicine(贝勒医学院)
The Platonic Universe: Do Foundation Models See the Same Sky?
UniverseTBD, :, Kshitij Duraphe, Michael J. Smith, Shashwat Sourav, John F. Wu
机构
*
Independent Researcher(独立研究者)
;
AstroAI
;
Harvard-Smithsonian CfA(哈佛-史密松天体物理中心)
;
University of Hertfordshire(赫特福德郡大学)
;
Washington University St. Louis(圣路易斯华盛顿大学)
;
Space Telescope Science Institute(空间望远镜科学研究所)
;
Johns Hopkins University(约翰霍普金斯大学)