Elmes*: Automated Construction of Fine-Grained Evaluation Rubrics for Large Language Models in Long-Tail Educational Scenarios
Elmes*:面向长尾教育场景的大语言模型细粒度评估量规自动构建
Tao Liu, Ye Lu, Ruohua Zhang, Siyu Song, Wentao Liu, Aimin Zhou, Hao Hao
机构
*
Shanghai Institute of AI for Education, East China Normal University(上海人工智能教育研究院,东华师范大学)
;
School of Computer Science and Technology, East China Normal University(计算机科学与技术学院,东华师范大学)
;
Shanghai Innovation Institute(上海创新研究院)
专题命中
评测与基准
:large language model(title,abstract);language model(title,abstract);LLM(abstract,abstract_cn);分类 cs.LG
Using Large Language Models to Support High Volume Application Review for an Undergraduate Research Program
使用大型语言模型支持本科研究项目的高容量申请评审
Varun Aggarwal, Kay Kobak, John Howarter
机构
*
Engineering Undergraduate Research Office, Purdue University(普渡大学本科生研究办公室)
;
Elmore School of Electrical and Computer Engineering, Purdue University(普渡大学电子与计算机工程学院)
;
School of Materials Engineering, Purdue University(材料工程学院)
专题命中
评测与基准
:large language model(title,abstract);language model(title,abstract);LLM(abstract,abstract_cn);分类 cs.CL
机构
*
East China University of Science and Technology, Shanghai, China(东华大学)
;
Renji Hospital Affiliated to Shanghai Jiaotong University School of Medicine, Shanghai, China(复旦大学附属中山医院)
专题命中
评测与基准
:large language model(title,abstract);language model(title,abstract);LLM(abstract,abstract_cn);分类 cs.AI
CommentsThis paper has been accepted by AAAI 2026. We update it for adding new evaluation results for ArxivRollBench-2025a and ArxivRollBench-2026a, with the evaluation of timly models like DeepSeekV4Pro, GPT-5.5, Claude-Opus-4.7, and so on. Source code: https://github.com/liangzid/ArxivRoll/ Online Leaderboard Website: https://arxivroll.moreoverai.com/
AgenticEval: Toward Agentic and Self-Evolving Safety Evaluation of Large Language Models
AgenticEval: 向大型语言模型的代理和自演化安全评估迈进
Yixu Wang, Xin Wang, Yang Yao, Xinyuan Li, Xibang Yang, Yan Teng, Xingjun Ma, Yingchun Wang
机构
*
Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
;
Fudan University(复旦大学)
;
The University of Hong Kong(香港大学)
;
East China Normal University(华东师范大学)
专题命中
评测与基准
:large language model(title,abstract);language model(title,abstract);LLM(summary_cn);分类 cs.AI
Yahan Li, Jifan Yao, John Bosco S. Bunyi, Adam C. Frank, Angel Hsing-Chi Hwang, Ruishan Liu
机构
*
Department of Computer Science, University of Southern California(南加州大学计算机科学系)
;
Department of Electrical and Computer Engineering, University of Southern California(南加州大学电气与计算机工程系)
;
Suzanne Dworak-Peck School of Social Work, University of Southern California(南加州大学苏兹安·德沃拉克-佩克社会工作学院)
;
Department of Psychiatry and the Behavioral Sciences, University of Southern California(南加州大学精神病学与行为科学系)
;
Annenberg School for Communication, University of Southern California(南加州大学安纳伯格通信学院)
专题命中
评测与基准
:large language model(title,abstract);language model(title,abstract);LLM(abstract,abstract_cn);分类 cs.CL
机构
*
Research Intern, Department of Mechanical and Aerospace Engineering, George Washington University(乔治华盛顿大学机械与航空航天工程系研究实习生)
;
Ph.D. Student, Department of Mechanical and Aerospace Engineering, George Washington University(乔治华盛顿大学机械与航空航天工程系博士生)
;
Undergraduate Student, Aerospace Program, University of California, Berkeley(加州大学伯克利分校航空航天项目本科生)
;
Full Professor, Department of Electrical Engineering and Computer Science, University of California, Berkeley(加州大学伯克利分校电气工程与计算机科学系教授)
;
Ph.D. Student, Department of Computer Science, George Washington University(乔治华盛顿大学计算机科学系博士生)
;
Associate Professor, Department of Mechanical and Aerospace Engineering, George Washington University(乔治华盛顿大学机械与航空航天工程系副教授)
专题命中
评测与基准
:large language model(title,abstract);language model(title,abstract);LLM(abstract,abstract_cn);分类 cs.AI