MedAraBench: Large-Scale Arabic Medical Question Answering Dataset and Benchmark
MedAraBench:大规模阿拉伯语医学问答数据集和基准测试
Mouath Abu-Daoud, Leen Kharouf, Omar El Hajj, Dana El Samad, Mariam Al-Omari, Jihad Mallat, Khaled Saleh, Nizar Habash, Farah E. Shamout
机构
*
Engineering Division, New York University Abu Dhabi(纽约大学阿布扎克分校工程系)
;
Cleveland Clinic Abu Dhabi(阿布扎克克利夫兰诊所)
;
Science Division, New York University Abu Dhabi(纽约大学阿布扎克分校科学系)
专题命中
评测与基准
:LLM(abstract);large language model(abstract);language model(abstract);分类 cs.CL
机构
*
State Key Lab of AI Safety, Institute of Computing Technology, CAS(人工智能安全国家重点实验室,计算技术研究所,中国科学院)
;
Key Lab of AI Safety, Chinese Academy of Sciences(人工智能安全重点实验室,中国科学院)
;
University of Chinese Academy of Sciences(中国科学院大学)
专题命中
评测与基准
:LLM(abstract);large language model(abstract);language model(abstract);分类 cs.CL
From Detection to Prevention: Explaining Security-Critical Code to Avoid Vulnerabilities
从检测到预防:解释安全关键代码以避免漏洞
Ranjith Krishnamurthy, Oshando Johnson, Goran Piskachev, Eric Bodden
机构
*
Paderborn University \& Fraunhofer IEM Paderborn Germany
;
Amazon Web Services Berlin Germany
;
Paderborn University \& Fraunhofer IEM
;
Amazon Web Services
专题命中
评测与基准
:LLM(abstract);large language model(abstract);language model(abstract);分类 cs.AI
机构
*
Department of Computer Science and Information Engineering, National Taiwan University, Taiwan(国家台湾大学计算机科学与信息工程系)
;
Institute of Information Science, Academia Sinica, Taiwan(台湾学术院信息科学研究所)
;
AI Research Center (AINTU), National Taiwan University, Taiwan(国家台湾大学人工智能研究中心(AINTU))
Reassessing Active Learning Adoption in Contemporary NLP: A Community Survey
重新评估当代自然语言处理中主动学习的采用情况:一项社区调查
Julia Romberg, Christopher Schröder, Julius Gonsior, Katrin Tomanek, Fredrik Olsson
机构
*
GESIS – Leibniz Institute for the Social Sciences(莱比锡社会科学研究所)
;
Center for Scalable Data Analytics and Artificial Intelligence (ScaDS.AI)(可扩展数据与人工智能研究中心)
;
Dresden/Leipzig, Leipzig University(莱比锡大学)
;
TUD Dresden University of Technology(德累斯顿技术大学)
;
All Ears
专题命中
评测与基准
:large language model(abstract);language model(abstract);分类 cs.CL、cs.LG
HalluHard: A Hard Multi-Turn Hallucination Benchmark
HalluHard: 一种具有挑战性的多轮 hallucination 评估基准
Dongyang Fan, Sebastien Delsad, Nicolas Flammarion, Maksym Andriushchenko
机构
*
EPFL(苏黎世联邦理工学院)
;
ELLIS Institute Tübingen(图宾根ELLIS研究所)
;
Max Planck Institute for Intelligent Systems(马克斯·普朗克智能系统研究所)
;
Tübingen AI Center(图宾根人工智能中心)
专题命中
评测与基准
:large language model(abstract);language model(abstract);分类 cs.CL、cs.AI
A Survey of AI Methods for Geometry Preparation and Mesh Generation in Engineering Simulation
工程仿真中几何准备和网格生成的AI方法综述
Steven Owen, Nathan Brown, Nikos Chrisochoides, Rao Garimella, Xianfeng Gu, Franck Ledoux, Na Lei, Roshan Quadros, Navamita Ray, Nicolas Winovich, Yongjie Jessica Zhang
机构
*
Sandia National Laboratories(桑迪亚国家实验室)
;
Old Dominion University(旧 Dominion 大学)
;
Los Alamos National Laboratory(洛斯阿拉莫斯国家实验室)
;
New York University / Stony Brook University(纽约大学 / 斯通布鲁克大学)
;
CEA(法国原子能委员会)
;
Dalian University of Technology(大连理工大学)
;
Carnegie Mellon University(卡内基梅隆大学)
专题命中
评测与基准
:large language model(abstract);language model(abstract);分类 cs.AI、cs.LG
机构
*
Mohamed bin Zayed University of Artificial Intelligence (MBZUAI)(莫扎伊德大学人工智能学院)
;
Ubiquitous Knowledge Processing Lab (UKP Lab)(通用知识处理实验室)
;
Department of Computer Science, TU Darmstadt(计算机科学系,图恩大学)
;
National Research Center for Applied Cybersecurity ATHENE(应用网络安全国家研究中心ATHENE)
专题命中
评测与基准
:large language model(abstract);language model(abstract);分类 cs.LG
机构
*
Department of Computer Science and Technology, Beijing Jiaotong University, Beijing, China(北京交通大学计算机科学与技术学院)
;
Tianjin Tasly Digital Chinese Medicine Technology Co., Ltd.(天津塔斯丽数字中医科技有限公司)
;
Tasly Biopharmaceuticals Co., Ltd.(塔斯丽生物医药有限公司)
;
State Key Laboratory of Chinese Medicine Modernization, Tianjin, China(中药现代化国家工程实验室,天津,中国)
;
Institute of Liver Diseases, Hubei Key Laboratory of the theory and application research of liver and kidney in traditional Chinese medicine, Hubei Provincial Hospital of Traditional Chinese Medicine, Wuhan, China(肝病研究所,湖北省中医肝肾理论与应用研究重点实验室,湖北省中医药研究院,武汉,中国)
;
Affiliated Hospital of Hubei University of Chinese Medicine, Wuhan, China(湖北中医药大学附属医院,武汉,中国)
;
Hubei Province Academy of Traditional Chinese Medicine, Wuhan, China(湖北省中医药研究院,武汉,中国)
;
Department of Gastroenterology, Guang’anmen Hospital, China Academy of Chinese Medical Sciences, Beijing, China(消化内科,广安门医院,中国中医科学院,北京,中国)
;
China Academy of Chinese Medical Sciences, Beijing, China(中国中医科学院,北京,中国)
;
China Institute for History of Medicine and Medical Literature, China Academy of Chinese Medical Sciences, Beijing, China(中国中医科学院中国医学史与医学文献研究所,北京,中国)
;
Beijing Research Institute of Chinese Medicine, Beijing University of Chinese Medicine, Beijing, China(北京中医研究院,北京中医药大学,北京,中国)
;
Beijing University of Chinese Medicine Third Affiliated Hospital, Beijing University of Chinese Medicine, Beijing 100029, China(北京中医药大学第三附属医院,北京中医药大学,北京100029,中国)
专题命中
评测与基准
:large language model(abstract);language model(abstract);分类 cs.AI
GEO-Bench-2: From Performance to Capability, Rethinking Evaluation in Geospatial AI
GEO-Bench-2:从性能到能力,重新思考地理空间AI的评估
Naomi Simumba, Nils Lehmann, Paolo Fraccaro, Hamed Alemohammad, Geeth De Mel, Salman Khan, Manil Maskey, Nicolas Longepe, Xiao Xiang Zhu, Hannah Kerner, Juan Bernabe-Moreno, Alexandre Lacoste
机构
*
IBM Research Europe(IBM欧洲研究院)
;
Technical University Munich(慕尼黑技术大学)
;
Clark University(克拉克大学)
;
MBZUAI
;
NASA Impact(NASA影响计划)
;
ESA Φ \Phi -lab(欧洲航天局Φ实验室)
;
Arizona State University(亚利桑那州立大学)
;
ServiceNow AI Research(ServiceNow人工智能研究)