Invariance Makes LLM Unlearning Resilient Even to Unanticipated Downstream Fine-Tuning
机构 * Michigan State University(密歇根州立大学) ; IBM Research(IBM研究院)
专题命中 安全评测 :safety(abstract);分类 cs.LG
Comments Accepted by ICML 2025
AI 大模型
大模型对齐、安全、越狱、红队、提示注入和可信评测。
机构 * Michigan State University(密歇根州立大学) ; IBM Research(IBM研究院)
专题命中 安全评测 :safety(abstract);分类 cs.LG
Comments Accepted by ICML 2025
机构 * Department of Mechanical Engineering, Tsinghua University(清华大学机械工程系) ; Department of Automation, Tsinghua University(清华大学自动化系)
专题命中 安全评测 :safety(abstract)
Comments Accepted by ITSC 2025
机构 * Department of Computer Science(计算机科学系) ; Cranberry-Lemon University(Cranberry-Lemon 大学) ; Texas A&M University(德克萨斯A&M大学) ; eBay Inc.(eBay公司)
专题命中 安全评测 :alignment(abstract)
专题命中 安全评测 :trustworthy(abstract)
专题命中 安全评测 :alignment(abstract)
Comments Accepted by IEEE 25th BIBE
机构 * Cornell University(康奈尔大学) ; Princeton University(普林斯顿大学) ; University of Washington(华盛顿大学)
专题命中 AI治理与伦理 :safety(abstract);分类 cs.CL、cs.AI、cs.CY
Comments 28 pages, 11 figures, 16 tables. In submission
机构 * Technical University of Munich(慕尼黑技术大学) ; LMU Munich(慕尼黑大学) ; Munich Center for Machine Learning(慕尼黑机器学习中心)
专题命中 AI治理与伦理 :alignment(abstract);safety(abstract);分类 cs.CL
Comments EMNLP 2025 (Oral)
专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI、cs.CY
机构 * Imperial College London(帝国理工学院) ; King's College London(国王学院) ; Independent Researcher(独立研究者)
专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI
Comments 12 pages,6 figures,Workshop on Technical AI Governance at ICML
专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI
专题命中 AI治理与伦理 :alignment(abstract)
Comments Accepted and Published in SBP-BRiMS 2025. 18th International Conference on Social Computing, Behavioral-Cultural Modeling & Prediction and Behavior Representation in Modeling and Simulation
机构 * The Chinese University of Hong Kong(香港中文大学) ; Fudan University(复旦大学)
专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI
机构 * University of Maryland, College Park(马里兰大学学院公园分校) ; Shandong Jiaotong University(山东交通大学)
专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.LG
机构 * Univ Gustave Eiffel(法国埃菲尔大学)
专题命中 其他安全 :safety(title,abstract);分类 cs.AI、cs.LG
Comments 18 pages, 5 figures, Accepted for publication in the proceedings of the 8th International Symposium on AI Verification SAIV 2025
专题命中 其他安全 :safety(title,abstract);分类 cs.CY
Comments Accepted to Symposium on Model Accountability, Sustainability and Healthcare (SMASH) 2025
机构 * Center for Information and Language Processing, LMU Munich(信息与语言处理中心,慕尼黑大学) ; Munich Center for Machine Learning (MCML)(慕尼黑机器学习中心(MCML)) ; Sorbonne Université, CNRS, ISIR, France(索邦大学,CNRS,ISIR,法国)
专题命中 其他安全 :alignment(title,abstract);分类 cs.CL
Comments preprint
专题命中 其他安全 :alignment(title,abstract);分类 cs.CL
Comments need updates
机构 * Shenzhen Research Institute of Big Data(大数据研究 institute) ; National Health Data Institute(国家健康数据研究所) ; Halmstad University(哈马碧大学) ; Shenzhen University of Advanced Technology(深圳先进技术大学)
专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.CL
机构 * DAMO Academy, Alibaba Group(达摩院,阿里巴巴集团)
专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI
Comments NeurIPS 2025
机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) ; Hong Kong University of Science and Technology(香港科学与技术大学)
专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI
机构 * Konkuk University(韩国康康大学)
专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI
Comments Accepted to EMNLP 2025 (Findings)
机构 * Robotics Institute, Carnegie Mellon University(卡内基梅隆大学机器人研究所) ; Machine Learning Department, Carnegie Mellon University(卡内基梅隆大学机器学习系) ; Department of Psychology, University of Pennsylvania(宾夕法尼亚大学心理学系) ; Neuroscience Institute, Carnegie Mellon University(卡内基梅隆大学神经科学研究所)
专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG
Comments 10 pages, 8 figures, 7 tables, NeurIPS 2025 Camera Ready Version (oral)
机构 * Department of Computer Science, ETH Zurich(苏黎世联邦理工学院计算机科学系) ; ETH AI Center(ETH人工智能中心) ; Ubiquitous Knowledge Processing Lab (UKP Lab), Department of Computer Science, Technical University of Darmstadt and National Research Center for Applied Cybersecurity ATHENE, Germany(达姆施塔特技术大学计算机科学系、应用网络安全国家研究中心ATHENE)
专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI
Comments Accepted to EMNLP 2025 Main as an oral presentation. David Dinucu-Jianu and Jakub Macina contributed equally. Code available: https://github.com/eth-lre/PedagogicalRL
机构 * Samsung Semiconductor, Inc.(三星半导体公司)
专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG
机构 * Michigan State University(密歇根州立大学) ; IBM Research(IBM研究院)
专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG
Comments Accepted by EMNLP 2025
机构 * MIT CSAIL(麻省理工学院计算机科学与人工智能实验室)
专题命中 其他安全 :alignment(abstract);分类 cs.LG
机构 * Meta Reality Labs(Meta现实实验室) ; Worcester Polytechnic Institute(沃斯特理工学院)
专题命中 其他安全 :alignment(abstract);分类 cs.CL
机构 * Southern University of Science and Technology(南方科技大学) ; The Chinese University of Hong Kong(香港中文大学)
专题命中 其他安全 :alignment(abstract);分类 cs.AI
机构 * KAIST(韩国科学技术院) ; University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校) ; Carnegie Mellon University(卡内基梅隆大学)
专题命中 其他安全 :alignment(abstract);分类 cs.CL
Comments Accepted to EMNLP 2025 Main Conference
机构 * SEACrowd ; Kreasof AI ; Universitas Indonesia(印度尼西亚大学) ; MBZUAI ; Capital One ; AI Singapore(AI新加坡) ; Cohere
专题命中 其他安全 :alignment(abstract);分类 cs.CL