机构
*
School of Artificial Intelligence, Beijing Normal University(北京师范大学人工智能学院)
;
Beijing Key Laboratory of Artificial Intelligence for Education(北京人工智能教育重点实验室)
;
Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所)
;
School of Cyber Security, University of Chinese Academy of Sciences(中国科学院大学网络安全学院)
;
Department of Computer Science and Technology, Institute for AI, Tsinghua University(清华大学人工智能研究院计算机科学与技术系)
;
Northeastern University(东北大学)
机构
*
University of Pittsburgh(匹兹堡大学)
;
Johns Hopkins University(约翰霍普金斯大学)
;
University of Notre Dame(诺特丹大学)
;
University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校)
;
University of Washington(华盛顿大学)
;
Allen Institute for Artificial Intelligence(人工智能研究院)
;
University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校)
专题命中
评测与基准
:LLM(summary_cn,abstract);large language model(title,abstract);language model(title,abstract);post-training(abstract)
Comments19 pages, 5 figures, 3 tables. Benchmark paper introducing SPLIT for evaluating empathy, linguistic naturalness, and cultural grounding in English and Ukrainian LLM responses
OntoLearner: A Modular Python Library for Ontology Learning with Large Language Models
OntoLearner: 一个用于大语言模型本体学习的模块化Python库
Hamed Babaei Giglou, Jennifer D'Souza, Andrei Aioanei, Nandana Mihindukulasooriya, Sören Auer
机构
*
TIB – Leibniz Information Centre for Science and Technology(TIB – 莱布尼茨科学与技术信息中心)
;
L3S Research Center, Leibniz University of Hannover(L3S研究中心,莱布尼茨汉诺威大学)
;
IBM Research(IBM研究院)
专题命中
评测与基准
:LLM(summary_cn,abstract);large language model(title,abstract);language model(title,abstract);分类 cs.AI
A Unified Framework for the Evaluation of LLM Agentic Capabilities
LLM 代理能力评估的统一框架
Pengyu Zhu, Lijun Li, Yaxing Lyu, Qianxin Luo, Jingyi Yang, Yi Liu, Tingfeng Hui, Xinyu Yuan, Li Sun, Sen Su, Jing Shao
机构
*
Beijing University of Posts and Telecommunications(北京邮电大学)
;
Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
;
Chongqing University of Posts and Telecommunications(重庆邮电大学)
;
North China Electric Power University(华北电力大学)
Automated grading of Linux/bash examinations using large language models: a four-level cognitive taxonomy approach
使用大语言模型对Linux/bash考试进行自动评分:一种四层认知分类方法
Manuel Alonso-Carracedo, Ruben Fernandez-Boullon, Pedro Celard, Francisco J. Rodriguez-Martinez, Lorena Otero-Cerdeira
机构
*
Universidade de Vigo, Department of Computer Science, ESEI-Higher School of Computer Engineering(维戈大学计算机科学系ESEI高等计算机工程学院)
;
IFCAE-Institute for Research in Physics, Computing and Aerospace Science(IFCAE物理、计算与航空航天科学研究所)
专题命中
评测与基准
:large language model(title,abstract);language model(title,abstract);prompting(abstract);分类 cs.CL、cs.AI
EduArt: An educational-level benchmark for evaluating art history knowledge in large language models
EduArt:评估大型语言模型艺术史知识的教育级基准
Gianmarco Spinaci, Lukas Klic, Giovanni Colavizza
机构
*
University of Bologna(博洛尼亚大学)
;
Villa i Tatti – The Harvard University Center for Italian Renaissance Studies(哈佛大学意大利文艺复兴研究中心(I Tatti))
;
University of Copenhagen(哥本哈根大学)
专题命中
评测与基准
:large language model(title,abstract);language model(title,abstract);分类 cs.CL
Challenges and Recommendations for LLMs-as-a-Judge in Multilingual Settings and Low-Resource Languages
多语言环境和低资源语言中LLM-as-a-Judge的挑战与建议
A. Seza Doğruöz, Xixian Liao, Verena Blaschke, Jakob Prange, Senyu Li, David Ifeoluwa Adelani
机构
*
LT3, IDLab, Universiteit Gent, Barcelona Supercomputing Center, LMU Munich & Munich Center for Machine Learning, German Center for Addiction Research in Childhood and Adolescence, University Medical Center Hamburg-Eppendorf, Mila - Quebec AI Institute, McGill University, Canada CIFAR AI Chair(LT3、IDLab、根特大学、巴塞罗那超级计算中心、慕尼黑莱茵河大学及慕尼黑机器学习中心、德国成年期成瘾研究中心、汉堡埃彭多夫大学医学中心、魁北克人工智能研究所、麦吉尔大学、加拿大 CIFAR 人工智能主席)