The Pluralistic Moral Gap: Understanding Judgment and Value Differences between Humans and Large Language Models
机构 * EPFL(苏黎世联邦理工学院) ; Bocconi University(博科尼大学)
专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI
Comments 13 pages, 4 figures
AI 大模型
大模型对齐、安全、越狱、红队、提示注入和可信评测。
机构 * EPFL(苏黎世联邦理工学院) ; Bocconi University(博科尼大学)
专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI
Comments 13 pages, 4 figures
机构 * Department of Electronics Technology, University of the Basque Country (UPV/EHU)(电子技术系,巴斯克国家大学(UPV/EHU)) ; Department of Electricity and Electronics, University of the Basque Country (UPV/EHU)(电力与电子系,巴斯克国家大学(UPV/EHU))
专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG
机构 * CARIAD SE(CARIAD公司) ; Karlsruhe Institute of Technology(卡尔斯鲁厄理工学院) ; FZI Research Center for Information Technology(弗劳恩霍夫信息技术研究中心)
专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.AI
机构 * Berkeley(伯克利)
专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.AI
机构 * HeartVoice Medical Technology(HeartVoice医疗科技) ; Saw Swee Hock School of Public Health and Institute of Data Science(Saw Swee Hock公共卫生学院和数据科学研究所) ; National University of Singapore(新加坡国立大学) ; Department of Cardiology, Peking University People’s Hospital(北京大学人民医院心内科) ; College of Integrative Chinese and Western Medicine, Anhui University of Chinese Medicine(安徽中医药大学整合中西医学学院) ; National Institute of Health Data Science, Peking University(北京大学国家健康数据科学研究院) ; Institute for Artificial Intelligence, Peking University(北京大学人工智能研究院)
专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.LG
机构 * Centre for Quantitative Medicine and Duke-NUS AI + Medical Science Initiative, Duke-NUS Medical School(定量医学中心和杜克-国立新加坡大学AI+医学科学计划,杜克-国立新加坡大学医学院) ; Graduate School of Engineering, The University of Tokyo(东京大学工程研究生院) ; Department of Population Health Sciences, Weill Cornell Medicine(流行病学与公共卫生科学系,韦尔·科恩医学中心) ; Department of Surgery, Duke University School of Medicine(外科医学系,杜克大学医学学院) ; Centre for Quantitative Medicine, Duke-NUS AI + Medical Science Initiative and Programme in Health Services and Systems Research, Duke-NUS Medical School and NUS Artificial Intelligence Institute, National University of Singapore(定量医学中心和杜克-国立新加坡大学AI+医学科学计划及健康服务与系统研究计划,杜克-国立新加坡大学医学院和新加坡国立大学人工智能研究所)
专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG
机构 * UBC(不列颠哥伦比亚大学) ; Microsoft(微软) ; Vector Institute for AI(人工智能向量研究所) ; CIFAR AI Chair(卡尔·弗雷德里克人工智能主席)
专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI
Comments arXiv admin note: substantial text overlap with arXiv:2407.14506
机构 * department of Electrical and Electronics Engineering, Koc University(电子与电气工程系,科克大学) ; Computer, Electrical, and Mathematical Sciences and Engineering Division, King Abdullah University of Science and Technology(计算机、电气和数学科学与工程系,国王阿卜杜勒-阿齐兹大学)
专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.LG
机构 * Faculty of Graduate Education Institute, Department of Artificial Intelligence (Interdisciplinary), Bahçeşehir University, Turkey(研究生教育学院人工智能系(跨学科)巴塞希尔大学,土耳其) ; School of Computer Science, University College Dublin, Ireland(计算机科学学院都柏林大学学院,爱尔兰) ; Faculty of Engineering and Natural Sciences, Department of Mathematics, Bahçeşehir University, Turkey(工程与自然科学学院数学系巴塞希尔大学,土耳其)
专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.LG
Comments 5th international Conference on Modelling, Computation and Optimization in Information Systems and Management Sciences (MCO 2025), June 4-6, 2025, Metz, France
机构 * Ufonia Limited(乌菲尼亚有限公司) ; University of York(约克大学)
专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI
Comments 29 pages
专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.LG
机构 * University of California, Berkeley(加州大学伯克利分校)
专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.CY
专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI
Comments NeurIPS 2024 Workshop on Adaptive Foundation Models
机构 * University of Oxford(牛津大学)
专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG
机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI
专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG
机构 * IIT Bombay(印度理工学院博伊斯分校) ; IBM Research(IBM研究院)
专题命中 安全评测 :DPO(abstract);分类 cs.CL、cs.AI
Comments 10 Pages, 5 Figures
专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI
专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.AI
Comments Accepted in ICCV 2025 workshops
机构 * University of Delaware(德克萨斯大学) ; Southern Methodist University(南方 Methodist 大学) ; Nemours Children’s Health(Nemours 儿童健康)
专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.LG
Comments This paper has been accepted in Machine Learning for Health (ML4H) Symposium. Link: https://proceedings.mlr.press/v259/fayyaz25a.html
机构 * Hong Kong Baptist University(香港 Baptist 大学) ; Shanxi University(山西大学) ; Shanghai Institute for Advanced Study of Zhejiang University(浙江大学上海先进研究院) ; Hong Kong University of Science and Technology(香港科技大学) ; The University of Manchester(曼彻斯特大学)
专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.AI
Comments 13 pages. 7 figures
Journal ref This paper is accpeted by ACL2025(Main)
专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI
Comments Preprint
机构 * Birla Institute of Technology and Science, Pilani, Hyderabad, India(比拉理工学院和科学学院,海得拉巴,印度)
专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI
Comments 12 pages, 4 figures
机构 * Centre pour la Sécurité de l'IA (CeSIA)(人工智能安全研究中心) ; École Polytechnique Fédérale de Lausanne (EPFL)(洛桑联邦理工学院)
专题命中 安全评测 :jailbreak(abstract);分类 cs.CL、cs.AI
专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI
Comments 34 pages, 12 figures
专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI
机构 * AA LAB, MODULABS(AA实验室,MODULABS) ; Aiffel, MODULABS(Aiffel,MODULABS) ; Department of Statistics, Inha university(统计系,Inha大学)
专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.CY
Comments Accepted at ICML 2025 Workshop on Multi-Agent Systems in the Era of Foundation Models: Opportunities, Challenges and Futures (MAS)
机构 * The Ohio State University(俄亥俄州立大学) ; Imperial College London(伦敦帝国理工学院) ; The University of Hong Kong(香港大学) ; University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校)
专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.LG
Comments ACL 2025
机构 * Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院) ; Shenzhen Institute of Advanced Technology Chinese Academy of Sciences(中国科学院深圳先进技术研究院)
专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG
Comments 35 pages, 15 figures, submitted to ACM Computing Surveys
机构 * Rutgers The State University of New Jersey(新泽西州立大学拉特格斯分校) ; University of Chicago(芝加哥大学)
专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI
Comments Accepted to ACM SIGKDD Explorations Volume 27 Issue 1