International AI Safety Report 2025: First Key Update: Capabilities and Risk Implications
专题命中 安全评测 :safety(title,abstract);AI safety(title,abstract);分类 cs.CY
AI 大模型
大模型对齐、安全、越狱、红队、提示注入和可信评测。
专题命中 安全评测 :safety(title,abstract);AI safety(title,abstract);分类 cs.CY
机构 * FutureAGI Inc.(未来人工智能公司)
专题命中 安全评测 :trustworthy(title,abstract);safety(abstract);prompt injection(abstract);分类 cs.CL、cs.AI
机构 * Institute of Software Chinese Academy of Sciences Beijing China(中国科学院软件研究所) ; Key Laboratory of System Software (Chinese Academy of Sciences) and State Key Laboratory of Computer Science, Institute of Software, Chinese Academy of Sciences, Beijing, China(中国科学院系统软件重点实验室和计算机科学国家重点实验室) ; University of Chinese Academy of Sciences, Beijing, China(中国科学院大学) ; Northeast University China(东北大学) ; Institute of Ai For industries Nanjing China(人工智能产业研究院) ; Southwest University China(西南大学) ; Nanyang Technological University Singapore(新加坡南洋理工大学) ; Alibaba China(阿里巴巴(中国))
专题命中 安全评测 :trustworthy(title,abstract);分类 cs.CL、cs.AI
Comments Accepted at NeurIPS 2025
机构 * Fudan University(复旦大学) ; Shanghai Innovation Institute(上海创新研究院)
专题命中 安全评测 :safety(title,abstract);分类 cs.CL、cs.AI
机构 * Independent Researcher(独立研究者)
专题命中 安全评测 :alignment(abstract);trustworthy(abstract);分类 cs.CL、cs.AI、cs.LG
Comments 13 pages. Code and data are available at https://github.com/strongSoda/LITERAL-TO-LIBERAL
机构 * Tsinghua University(清华大学) ; ByteDance Seed(字节跳动种子) ; Princeton University(普林斯顿大学)
专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.AI
机构 * University of Alberta(阿尔伯塔大学) ; Mila - Quebec Artificial Intelligence Institute(魁北克人工智能研究所) ; The University of Tokyo(东京大学) ; Macau University of Science and Technology(澳门科学技术大学)
专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.AI
Comments 4 pages, 2 figures, To appear in ASE 2025 Demo Track
机构 * Princeton University(普林斯顿大学) ; CISPA Helmholtz Center for Information Security(CISPA海德堡信息安全中心) ; Michigan State University(密歇根州立大学) ; Ohio State University(俄亥俄州立大学) ; Brown University(布朗大学) ; Massachusetts Institute of Technology(麻省理工学院) ; University of California, Los Angeles(加州大学洛杉矶分校) ; University of Tübingen(图宾根大学) ; Old Dominion University(旧 Dominion 大学) ; Technical University of Munich(慕尼黑技术大学) ; Cornell University(康奈尔大学) ; Forest AI(森林AI)
专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG
Comments 14 pages; Accepted to NeurIPS 2025. Link to poster: https://neurips.cc/virtual/2025/poster/121919; Link to project website: https://www.peerbench.ai/
机构 * Georgia Institute of Technology(佐治亚理工学院) ; Purdue University(普渡大学)
专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.AI
机构 * University of Wisconsin-Madison(威斯康星大学麦迪逊分校) ; LG AI Research(LG人工智能研究) ; Seoul National University of Science and Technology(首尔科学技术大学) ; New York University(纽约大学) ; Genentech(基因泰克)
专题命中 安全评测 :alignment(abstract);分类 cs.LG
Comments 20 pages, 11 figures
机构 * LG Electronics USA(LG电子美国公司)
专题命中 安全评测 :alignment(abstract);分类 cs.CL
Comments Accepted to PMLR and NeurIPS 2025 UniReps
机构 * LARG ; Research Center for Social Computing and Interactive Robotics(社会计算与交互机器人研究室) ; Harbin Institute of Technology(哈尔滨工业大学) ; School of Computer Science and Engineering(计算机科学与工程学院) ; Central South University(中南大学) ; The University of Hong Kong(香港大学) ; ByteDance China (Seed)(字节跳动中国(种子))
专题命中 安全评测 :alignment(abstract);分类 cs.CL
Comments Preprint. Code: https://github.com/LightChen233/AutoPR . Benchmark: https://huggingface.co/datasets/yzweak/PRBench
机构 * Guangdong Provincial Key Laboratory of Intelligent Information Processing(广东省智能信息处理重点实验室) ; Shenzhen Key Laboratory of Media Security(深圳媒体安全重点实验室) ; SZU AFS Joint Innovation Center for AI Technology, Shenzhen University(深圳大学AFS人工智能技术联合创新中心) ; University of North Texas(北卡罗来纳州立大学) ; Academy of Forensic Science(法医科学研究院) ; Nanchang University(南昌大学)
专题命中 安全评测 :alignment(abstract)
机构 * School of Data Science, Fudan University(复旦大学数据科学学院) ; University of Surrey(萨里大学)
专题命中 安全评测 :safety(abstract)