A Survey on Evaluating Quality and Trustworthiness in LLM-Generated Data
评估LLM生成数据的质量与可信度综述
Kaituo Zhang, Mingzhi Hu, Hoang Anh Duy Le, Fariha Kabir Torsha, Zhimeng Jiang, Minh Khai Bui, Chia-Yuan Chang, Yu-Neng Chuang, Zhen Xiong, Ying Lin, Guanchu Wang, Na Zou
机构
*
University of Houston(德克萨斯大学休斯敦分校)
;
Worcester Polytechnic Institute(沃思利理工学院)
;
Rice University(里德大学)
;
Texas A&M University(德克萨斯农工大学)
;
University of Wisconsin - Madison(威斯康星大学麦迪逊分校)
;
University of Southern California(南加州大学)
;
University of North Carolina at Charlotte(北卡罗来纳州立大学夏洛特分校)
专题命中
评测与基准
:LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.AI、cs.LG
Analysing Differences in Persuasive Language in LLM-Generated Text: Uncovering Stereotypical Gender Patterns
分析LLM生成文本中说服性语言的差异:揭示刻板的性别模式
Amalie Brogaard Pauli, Maria Barrett, Max Müller-Eberstein, Isabelle Augenstein, Ira Assent
机构
*
Department of Computer Science, Aarhus University(阿arhus大学计算机科学系)
;
AMD Silo AI
;
University of Tokyo(东京大学)
;
IT University of Copenhagen(哥本哈根IT大学)
;
Department of Computer Science, University of Copenhagen(哥本哈根大学计算机科学系)
专题命中
评测与基准
:LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI
From Scoring to Explanations: Evaluating SHAP and LLM Rationales for Rubric-based Teaching Quality Assessment
从评分到解释:评估基于量规的教学质量评估中的SHAP和LLM理由
Ivo Bueno, Babette Bühler, Philipp Stark, Tim Fütterer, Ulrich Trautwein, Dorottya Demszky, Heather Hill, Enkelejda Kasneci
机构
*
Technical University of Munich(慕尼黑技术大学)
;
Munich Center for Machine Learning (MCML)(慕尼黑机器学习中心)
;
Lund University(吕勒奥大学)
;
University of Tübingen(图宾根大学)
;
Stanford Graduate School of Education(斯坦福大学教育研究生院)
;
Harvard Graduate School of Education(哈佛大学教育研究生院)
专题命中
评测与基准
:LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI
Consistent and Distinctive: LLM Benchmark Efficiency via Maximum Independent Set Prompt Selection on Similarity Graphs
一致且独特:基于相似图最大独立集提示选择的LLM基准测试效率
Denica Kjorvezir, Marko Djukanović, Ana Gjorgjevikj, Gjorgjina Cenikj, Tome Eftimov
机构
*
Computer Systems Department, Jožef Stefan Institute, Ljubljana, Slovenia(计算机系统部,乔塞夫·斯塔芬研究所,卢布尔雅那,斯洛文尼亚)
;
Jožef Stefan International Postgraduate School, Ljubljana, Slovenia(乔塞夫·斯塔芬国际研究生学院,卢布尔雅那,斯洛文尼亚)
;
Center for Astrophysics and Cosmology, University of Nova Gorica, Nova Gorica, Slovenia(天体物理与宇宙学中心,诺瓦戈里察大学,诺瓦戈里察,斯洛文尼亚)
专题命中
评测与基准
:LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI
Modeling Community Attitude through Reaction Tone: A Human-AI Collaborative Framework for Evaluating LLM Alignment with Linguistic Behaviors in Online Communities
通过反应语气建模社区态度:评估LLM与在线社区语言行为对齐的人机协作框架
Nuan Wen, Xuezhe Ma
机构
*
Information Sciences Institute University of Southern California(南加州大学信息科学研究所)
专题命中
评测与基准
:LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI
机构
*
Department of Medicinal Chemistry, Faculty of Pharmacy, Tehran University of Medical Sciences(药学系,泰赫兰医科大学)
;
Department of Computer Sciences, Faculty of Mathematics and Computer Sciences, Amir Kabir University of Technology(计算机科学系,阿米尔·卡比尔技术大学)
;
Department of Mathematical Sciences, Sharif University of Technology(数学科学系,沙菲克技术大学)
;
Department of Computer Sciences, Missouri University of Science and Technology(计算机科学系,密苏里科学与技术大学)
;
Department of Computer Engineering, Sharif University of Technology(计算机工程系,沙菲克技术大学)
;
Department of Faculty of Interdisciplinary Science and Technology, Tarbiat Modares University(跨学科科学与技术学院,塔里亚特莫达res大学)
;
Electronics Research Institute, Sharif University of Technology(电子研究所,沙菲克技术大学)
;
Department of Electrical Engineering, Sharif University of Technology(电气工程系,沙菲克技术大学)
;
The Alan Turing Institute, London, United Kingdom(艾伦·图灵研究所,伦敦,英国)
;
Department of Radiation Oncology, Massachusetts General Hospital & Harvard Medical School(放射肿瘤科,麻省总医院及哈佛医学院)
;
Health Informatics Lab, Metropolitan College, Boston University(健康信息学实验室,波士顿大学)
专题命中
评测与基准
:LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI
Comments14 pages, 2 figures, 2 tables. The revised version includes McNemar's paired statistical analysis, Wilson confidence intervals, expanded methodological clarifications, a revised discussion of evidence retrieval, improved reproducibility details, and updated limitations
A Scoping Review of LLM-as-a-Judge in Healthcare and the MedJUDGE Framework
对LLM-as-a-Judge在医疗领域的综述及MedJUDGE框架
Chenyu Li, Zohaib Akhtar, Mingu Kwak, Yuelyu Ji, Hang Zhang, Tracey Obi, Yufan Ren, Xizhi Wu, Sonish Sivarajkumar, Harold P. Lehmann, Shyam Visweswaran, Michael J. Becich, Danielle L. Mowery, Renxuan Liu, Haoyang Sun, Yanshan Wang
机构
*
Department of Biomedical Informatics, School of Medicine, University of Pittsburgh(匹兹堡大学医学院生物医学信息学系)
;
Department of Health Information Management, School of Health and Rehabilitation Sciences, University of Pittsburgh(匹兹堡大学健康与康复科学学院健康信息管理系)
;
OpenCura, Health Innovation Consortium(OpenCura健康创新联盟)
;
Northwestern University, Kellogg School of Management(西北大学凯洛格管理学院)
;
Intelligent Systems Program, School of Computing and Information, University of Pittsburgh(匹兹堡大学计算与信息学院智能系统项目)
;
Johns Hopkins University School of Medicine Biomedical Informatics and Data Science(约翰霍普金斯大学医学院生物医学信息学与数据科学)
;
Clinical and Translational Science Institute, University of Pittsburgh(匹兹堡大学临床与转化科学研究所)
;
Institute for Biomedical Informatics, University of Pennsylvania(宾夕法尼亚大学生物医学信息学研究所)
;
Data Science, School of Computing and Information, University of Pittsburgh(匹兹堡大学计算与信息学院数据科学)
专题命中
评测与基准
:LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI
Learning to Ask: When LLM Agents Meet Unclear Instruction
学习提问:当LLM代理遇见模糊指令
Wenxuan Wang, Juluan Shi, Zixuan Ling, Yuk-Kit Chan, Chaozheng Wang, Cheryl Lee, Youliang Yuan, Jen-tse Huang, Wenxiang Jiao, Michael R. Lyu
机构
*
Renmin University of China(中国人民大学)
;
The Chinese University of Hong Kong(香港中文大学)
;
The Chinese University of Hong Kong, Shenzhen(香港中文大学深圳校区)
;
Johns Hopkins University(约翰霍普金斯大学)
;
Xiaohongshu Inc.(小红书公司)
专题命中
评测与基准
:LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI