Fairness Evaluation of Large Language Models in Academic Library Reference Services
学术图书馆参考服务中大型语言模型的公平性评估
Haining Wang, Jason Clark, Yueru Yan, Star Bradley, Ruiyang Chen, Yiqiong Zhang, Hengyi Fu, Zuoyu Tian
机构
*
Indiana University(印第安纳大学)
;
Montana State University(蒙塔纳州立大学)
;
Wuhan University(武汉大学)
;
Guangdong University of Foreign Studies(广东外语外贸大学)
;
San José State University(圣何塞州立大学)
;
Macalester College(麦基尔学院)
专题命中
评测与基准
:large language model(title,abstract);language model(title,abstract);prompting(abstract);分类 cs.CL、cs.AI
机构
*
Center for Postgraduate Clinical Training and Career Development, Nagoya University Hospital(名古屋大学医院研究生临床培训与职业发展中心)
;
Center for Medical Education, Graduate School of Medicine, Nagoya University(名古屋大学医学研究生院医学教育中心)
;
Scientific Research Works Peer Support Group (SRWS-PSG)(科学研究作品同伴支持组)
;
Department of Internal Medicine, Kyoto Min-iren Asukai Hospital(京都Min-iren Asukai医院内科学部)
;
Department of Healthcare Epidemiology, Kyoto University Graduate School of Medicine / School of Public Health(京都大学医学研究院/公共卫生学院卫生流行病学部)
;
Department of International and Community Oral Health, Tohoku University Graduate School of Dentistry(东京东医齿学研究院国际与社区口腔健康部)
;
Department of Psychiatry, Okayama Psychiatric Medical Center(冈山精神医学中心精神病学部)
;
CureApp, Inc.(CureApp公司)
;
Department of Psychiatry, Seichiryo Hospital(Seichiryo医院精神病学部)
;
Oku medical clinic(Oku医疗诊所)
;
Department of Health Promotion and Human Behavior, Kyoto University Graduate School of Medicine / School of Public Health(京都大学医学研究院/公共卫生学院健康促进与人类行为部)
;
Kyoto University Hospital(京都大学医院)
;
Division of Radiology and Biomedical Engineering, Graduate School of Medicine, The University of Tokyo(东京大学医学研究生院放射学与生物医学工程部)
;
Department of Rehabilitation, Kurashiki Medical Centre(Kurashiki医疗中心康复部)
;
Department of Epidemiology, Graduate School of Medicine, Dentistry, and Pharmaceutical Sciences, Okayama University(冈山大学医学、牙科和药学研究生院流行病学部)
专题命中
评测与基准
:large language model(title,abstract);language model(title,abstract);LLM(abstract);分类 cs.AI
Comments9 pages, 1 figure, 4 tables. Presented at the 3rd International Workshop on Value Engineering in AI (VALE 2025), 28th European Conference on AI. To appear in Springer LNCS
ToolHaystack: Stress-Testing Tool-Augmented Language Models in Realistic Long-Term Interactions
ToolHaystack: 在现实长期交互中压力测试工具增强的语言模型
Beong-woo Kwak, Minju Kim, Dongha Lim, Hyungjoo Chae, Dongjin Kang, Sunghwan Kim, Dongil Yang, Jinyoung Yeo
机构
*
Department of Artificial Intelligence, Yonsei University(人工智能系,延世大学)
;
Department of Computer Science & Engineering, Yonsei University(计算机科学与工程系,延世大学)
专题命中
评测与基准
:language model(title,abstract);large language model(abstract);分类 cs.CL
SePer: Measure Retrieval Utility Through The Lens Of Semantic Perplexity Reduction
SePer:通过语义困惑度降低的视角衡量检索效用
Lu Dai, Yijie Xu, Jinhui Ye, Hao Liu, Hui Xiong
机构
*
The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州))
;
The Hong Kong University of Science and Technology(香港科学与技术大学)
;
Carnegie Mellon University(卡内基梅隆大学)
专题命中
评测与基准
:LLM(abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI、cs.LG
机构
*
Department of Mathematics and Computer Science University of Cagliari(数学与计算机科学系卡利亚里大学)
;
Department of Computer Science University of Salerno(计算机科学系萨勒诺大学)
专题命中
评测与基准
:large language model(abstract);language model(abstract);分类 cs.AI
CommentsThis is a post-peer-review, pre-copyedit version to be published in the Prooceedings of the 33rd Symposium On Advanced Database Systems (SEBD 2025), 7 pages, 4 figures
Fantastic Bugs and Where to Find Them in AI Benchmarks
人工智能基准中的非凡虫子及其寻找方法
Sang Truong, Yuheng Tu, Michael Hardy, Anka Reuel, Zeyu Tang, Jirayu Burapacheep, Jonathan Perera, Chibuike Uwakwe, Ben Domingue, Nick Haber, Sanmi Koyejo
Comments28 pages, 16 figures, this is an original manuscript of an article published by Taylor & Francis in the International Journal of Human-Computer Interaction (IJHCI), available online: https://doi.org/10.1080/10447318.2025.2594750
AeroVerse: UAV-Agent Benchmark Suite for Simulating, Pre-training, Finetuning, and Evaluating Aerospace Embodied World Models
AeroVerse:用于模拟、预训练、微调和评估航空航天具身世界模型的UAV-Agent基准套件
Fanglong Yao, Yuanchang Yue, Youzhi Liu, Xian Sun, Kun Fu
机构
*
Aerospace Information Research Institute, Chinese Academy of Sciences(中国科学院航空信息研究所)
;
University of Chinese Academy of Sciences(中国科学院大学)
;
School of Electronic, Electrical and Communication Engineering, University of Chinese Academy of Sciences(中国科学院大学电子电气与通信工程学院)
;
Key Laboratory of Target Cognition and Application Technology(TCAT), Aerospace Information Research Institute, Chinese Academy of Sciences(中国科学院航空信息研究所目标认知与应用技术重点实验室)
Peican Lin, Gan Sun, Chenxi Liu, Fazeng Li, Weihong Ren, Yang Cong
机构
*
School of Automation Science and Engineering, South China University of Technology(华南理工大学自动化科学与工程学院)
;
State Key Laboratory of Robotics, Shenyang Institute of Automation, Chinese Academy of Sciences(中国科学院沈阳自动化研究所机器人重点实验室)
;
State Key Laboratory of Robotics and System, School of Mechanical Engineering and Automation, Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳)机械工程与自动化学院机器人与系统重点实验室)
专题命中
评测与基准
:language model(abstract)
AI总结
OpenVLN通过强化学习和长视距规划器提升无人机在复杂空中环境中的长视距导航能力。
CommentsContent: 8 pages 4 figures, conference paper under review