Safety and Security Analysis of Large Language Models: Benchmarking Risk Profile and Harm Potential
专题命中 安全评测 :safety(title,abstract);alignment(abstract);分类 cs.CY
AI 大模型
大模型对齐、安全、越狱、红队、提示注入和可信评测。
专题命中 安全评测 :safety(title,abstract);alignment(abstract);分类 cs.CY
机构 * Massachusetts Institute of Technology, USA(麻省理工学院, 美国) ; Seoul National University Hospital, South Korea(首尔国立大学医院, 韩国) ; Doctor Diary, South Korea(Doctor Diary, 韩国) ; Google Research, USA(谷歌研究, 美国) ; Independent Researcher(独立研究者)
专题命中 安全评测 :safety(title,abstract);分类 cs.AI
Comments Accepted by Proceedings of Interspeech 2025; Website: https://han811.github.io/VocalAgent2025/
专题命中 安全评测 :safety(title,abstract);分类 cs.AI
机构 * CUHK MMLab(香港中文大学多媒体实验室) ; SenseTime Research(商汤科技研究院) ; CPII under InnoHK(创新香港下的CPII)
专题命中 安全评测 :alignment(abstract);safety(abstract);分类 cs.CL、cs.AI
专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG
专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.LG
Comments Accepted & presented at THE 16th INTERNATIONAL IEEE CONFERENCE ON COMPUTING, COMMUNICATION AND NETWORKING TECHNOLOGIES (ICCCNT) 2025
Journal ref THE 16th INTERNATIONAL IEEE CONFERENCE ON COMPUTING, COMMUNICATION AND NETWORKING TECHNOLOGIES (ICCCNT) 2025
机构 * Peking University(北京大学) ; National University of Singapore(新加坡国立大学) ; Institute of Science Tokyo(东京科学研究院) ; Nanjing University(南京大学) ; Westlake University(西湖大学) ; Southeast University(东南大学)
专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.AI
Comments 22 pages, 9 figures, 6 tables
专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.AI
Comments Accepted to EMNLP 2025; 15 pages; 5 figures, 11 tables; Code available at https://github.com/eunseongc/CARE
机构 * Luxembourg Institute of Science and Technology (LIST)(卢森堡科学与技术研究院) ; RMT Labs(RMT实验室)
专题命中 安全评测 :trustworthy(abstract);分类 cs.CL
Comments Accepted at the IEEE GLOBECOM Workshops 2025: "Large AI Model over Future Wireless Networks"
机构 * University of Southern California(南加州大学) ; University of Queensland(昆士兰大学) ; University of California, San Diego(加州大学圣地亚哥分校) ; University of Buffalo(布法罗大学) ; University of California, Merced(加州大学默塞德分校)
专题命中 安全评测 :alignment(abstract);分类 cs.CL
机构 * Amazon(亚马逊公司) ; Alexa, Amazon(亚马逊Alexa部门) ; Virginia Tech(弗吉尼亚理工大学) ; University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
专题命中 安全评测 :alignment(abstract);分类 cs.LG
Comments 17 pages, 9 figures, 14 tables, Findings of the Association for Computational Linguistics: EMNLP 2025
机构 * Youngstown State University(扬斯敦州立大学)
专题命中 安全评测 :alignment(abstract);分类 cs.CL
机构 * Ixent Games
专题命中 安全评测 :trustworthy(abstract);分类 cs.AI
Comments Version 2: This article consolidates and replaces a previous version to present the complete research in a single, comprehensive manuscript
机构 * Sun Yat-sen University(中山大学)
专题命中 安全评测 :safety(abstract)
Comments 13 pages, 6 figures
机构 * Nanjing University of Science and Technology(南京理工大学) ; Intellifusion Inc.(Intellifusion公司) ; Northwestern Polytechnical University(西北工业大学) ; Zhejiang Lab(浙江实验室) ; Yan’an University(延安大学) ; Beijing University of Posts and Telecommunications(北京邮电大学)
专题命中 安全评测 :alignment(abstract)