机构
*
The International Joint Institute of Tianjin University(天津大学国际联合研究院)
;
Tianjin University(天津大学)
;
TJUNLP Lab, School of Computer Science and Technology, Tianjin University(天津大学计算机科学与技术学院)
;
National Governance Institute, Tianjin Normal University(天津师范大学国家治理研究院)
;
College of Computer and Information Engineering, Tianjin Normal University(天津师范大学计算机与信息工程学院)
专题命中
评测与基准
:large language model(title,abstract);language model(title,abstract);LLM(abstract,abstract_cn);分类 cs.CL、cs.AI
Comments69 pages (main text + four appendices: A. Mathematical Theory and Derivations B. generation/evaluation prompts, C. council-evaluation excerpt D. Auditing the Metanym Game-GPQA correlation: No Leaks Found), 3 figures, 19 tables. Github repo with pages/figures/tables/code and data for reproducing results: this https URL (https://github.com/dnordfors/metanym-game-paper)
Designing a Robust LLM-Based Evaluation System for Agentic AI in Drug Discovery Through Human Alignment
通过人类对齐设计用于药物发现中智能体AI的鲁棒性大语言模型评估系统
Emma Granqvist, Rocío Mercado, Samuel Genheden
机构
*
AstraZeneca(阿斯利康)
;
Chalmers University of Technology(查尔姆斯理工大学)
;
University of Gothenburg(哥德堡大学)
;
Science for Life Laboratory (SciLifeLab)(生命科学实验室)
专题命中
评测与基准
:LLM(title,summary_cn);large language model(abstract);language model(abstract);分类 cs.LG
Comments37 pages, 8 figures. Published in Transactions on Machine Learning Research (TMLR), 2026. Supplementary material included as ancillary material
No PUN Intended: Plausible Unknown Names for Person-Centred LLM Evaluation
并非双关:面向以人为中心的大语言模型评估的合理未知姓名(PUN)
Dimitri Staufer, David Hartmann, Ibrahim Baroud
机构
*
Technische Universität Berlin(柏林工业大学)
;
Weizenbaum Institute for the Networked Society(魏茨曼网络社会研究所)
;
German Research Center for Artificial Intelligence (DFKI)(德国人工智能研究中心)