Cross-Modal Safety Alignment: Is textual unlearning all you need?
机构 * University of California, Riverside(加州大学河滨分校)
专题命中 偏好对齐 :alignment(title,abstract);safety(title,abstract);RLHF(abstract);分类 cs.CL、cs.LG
Comments Accepted by EMNLP 2024 Findings
AI 大模型
大模型对齐、安全、越狱、红队、提示注入和可信评测。
机构 * University of California, Riverside(加州大学河滨分校)
专题命中 偏好对齐 :alignment(title,abstract);safety(title,abstract);RLHF(abstract);分类 cs.CL、cs.LG
Comments Accepted by EMNLP 2024 Findings
机构 * The Chinese University of Hong Kong(香港中文大学) ; Fudan University(复旦大学)
专题命中 偏好对齐 :alignment(title,abstract);DPO(abstract);分类 cs.CL、cs.AI
专题命中 偏好对齐 :DPO(title,abstract);alignment(abstract);分类 cs.CL
专题命中 偏好对齐 :alignment(title,abstract);分类 cs.CL、cs.AI
Comments NeurIPS 2025
机构 * Yale University(耶鲁大学) ; Allen Institute for AI(人工智能研究院)
专题命中 偏好对齐 :alignment(abstract);RLHF(abstract);分类 cs.CL、cs.AI、cs.LG
机构 * Sun Yat-sen University(中山大学) ; X-Era AI Lab(X-Era人工智能实验室)
专题命中 偏好对齐 :alignment(abstract);DPO(abstract);分类 cs.AI、cs.LG
机构 * University of Cambridge(剑桥大学) ; Apta
专题命中 偏好对齐 :alignment(abstract);分类 cs.LG
专题命中 偏好对齐 :RLHF(abstract);分类 cs.CL
Comments 5 Pages, 4 Figures, 4 Tables
Journal ref 39th Conference on Neural Information Processing Systems, 2025, Workshop: Reliable ML from Unreliable Data
机构 * Alibaba AAIG(阿里巴巴AAIG)
专题命中 安全训练 :alignment(title,abstract);safety(title,abstract);jailbreak(abstract);分类 cs.CL、cs.AI、cs.CY
Comments Technical Report Code & Model weights available: https://github.com/Alibaba-AAIG/Oyster
专题命中 安全训练 :safety(title,abstract);分类 cs.CL、cs.AI
Comments Main Text: 2943; Abstract: 256; Tables and Figures: 5
机构 * CASIA(中国科学院自动化研究所) ; ByteDance Seed(字节跳动种子实验室) ; UCAS(中国科学院大学) ; FiveAges ; NJU(南京大学)
专题命中 安全训练 :alignment(title,abstract);分类 cs.AI
Comments NeurIPS 2025
机构 * University of Wisconsin-Madison(威斯康星大学麦迪逊分校) ; Tsinghua University(清华大学)
专题命中 安全训练 :alignment(abstract);safety(abstract);分类 cs.CL
专题命中 安全训练 :safety(abstract);分类 cs.AI
机构 * Department of Mechanical Engineering, Vrije Universiteit Brussel(布鲁塞尔自由大学机械工程系) ; imec ; Flanders Make(弗拉芒制造) ; Artificial Intelligence (AI) Lab, Vrije Universiteit Brussel(布鲁塞尔自由大学人工智能实验室)
专题命中 安全训练 :safety(abstract)
机构 * The Chinese University of Hong Kong(香港中文大学) ; Noah’s Ark Lab, Huawei(华为诺亚实验室) ; City University of Hong Kong(香港城市大学)
专题命中 越狱攻击 :jailbreak(title,abstract);safety(abstract);分类 cs.CL、cs.LG
机构 * Arizona State University(亚利桑那州立大学)
专题命中 越狱攻击 :safety(abstract);分类 cs.CL、cs.AI
Comments Published in Reliable ML from Unreliable Data workshop @ NeurIPS 2025
机构 * Independent Researcher(独立研究者)
专题命中 越狱攻击 :prompt injection(abstract,comments);分类 cs.AI
Comments 33 pages, 3 figures, 6 tables. Keywords: LLM security; defense-in-depth; prompt injection; activation steering; multimodal sandbox; threat modeling
机构 * Databricks
专题命中 红队测试 :red teaming(title,abstract);safety(abstract);分类 cs.AI
专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.AI
Comments The paper is not yet mature and needs further improvement
机构 * ENSIAS, Mohammed V University in Rabat(ENSIAS,摩洛哥拉巴特穆罕默德五世大学) ; Samovar, Télécom SudParis, Institut Polytechnique de Paris(Samovar,电信南巴黎,巴黎理工学院)
专题命中 隐私与版权 :alignment(abstract);分类 cs.AI、cs.LG
Comments 6 pages, 3 figures, conference
机构 * University of Oxford(牛津大学) ; FLAIR University of Oxford(牛津大学FLAIR)
专题命中 安全评测 :safety(title,abstract);AI safety(title);分类 cs.AI
机构 * Hong Kong University of Science and Technology(香港理工大学) ; Peking University(北京大学) ; University of Edinburgh(爱丁堡大学)
专题命中 安全评测 :safety(title,abstract);jailbreak(abstract);分类 cs.CL、cs.AI
机构 * University of Quebec at Chicoutimi(魁北克大学恰普蒂米分校)
专题命中 安全评测 :safety(abstract);分类 cs.AI
机构 * ARTORG Center for Biomedical Engineering Research, University of Bern(ARTORG生物医学工程研究中心,伯恩大学) ; Inselspital (Bern University Hospital)(Inselspital(伯恩大学医院)) ; Insel Gruppe Bern Universitätsinstitut für Diagnostische, Interventionelle und Pädiatrische Radiologie(Bern大学诊断、介入和儿科放射学研究所) ; Department of Radiation Oncology, Inselspital, Bern University Hospital(放射肿瘤科,Inselspital,伯恩大学医院)
专题命中 安全评测 :alignment(abstract);分类 cs.AI
Comments Accepted by iMIMIC at MICCAI 2025
机构 * Université Côte d’Azur, CNRS, Inria, I3S, France(法国大学-科蒂-阿祖尔大学、国家科学研究中心、法国国家信息与自动化研究所、I3S研究所)
专题命中 安全评测 :trustworthy(abstract);分类 cs.CL
机构 * University of California, Merced(加州大学梅尔德分校) ; vivo Mobile Communication Co., Ltd.(vivo移动通信有限公司) ; University of Queensland(昆士兰大学) ; UCLA(加州大学洛杉矶分校) ; University at Buffalo(布法罗大学)
专题命中 安全评测 :trustworthy(abstract);分类 cs.CL
机构 * Shanghai Jiao Tong University(上海交通大学)
专题命中 安全评测 :safety(abstract);分类 cs.AI
机构 * School of Information Management & Engineering, Shanghai University of Finance and Economics(上海金融学院信息管理与工程学院) ; Key Laboratory of Data Intelligence and Management (Beihang University), Ministry of Industry and Information Technology, School of Economics and Management, Beihang University(北京航空航天大学数据智能与管理重点实验室)
专题命中 安全评测 :trustworthy(abstract);分类 cs.LG
Comments 42 pages
机构 * Turing Inc.(图灵公司)
专题命中 安全评测 :safety(abstract)
Comments 2nd Place Winner, ICCV 2025 2COOOL Competition
机构 * Tsinghua University(清华大学)
专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL
Comments NeurIPS 2025 Spotlight