Improving LLM Safety Alignment with Dual-Objective Optimization
机构 * University of California, Berkeley(加州大学伯克利分校)
专题命中 偏好对齐 :alignment(title,abstract);safety(title,abstract);DPO(abstract);jailbreak(abstract)
Comments ICML 2025
AI 大模型
大模型对齐、安全、越狱、红队、提示注入和可信评测。
机构 * University of California, Berkeley(加州大学伯克利分校)
专题命中 偏好对齐 :alignment(title,abstract);safety(title,abstract);DPO(abstract);jailbreak(abstract)
Comments ICML 2025
机构 * BAISH | UBA | Apart Research(BAISH | UBA | Apart研究) ; University of São Paulo(圣保罗大学) ; Apart Research(Apart研究) ; Dovetail Research | Apart Research(Dovetail研究 | Apart研究)
专题命中 偏好对齐 :RLHF(abstract);分类 cs.CL、cs.AI、cs.LG
机构 * Stanford University(斯坦福大学) ; University of Toronto(多伦多大学) ; University of Pennsylvania(宾夕法尼亚大学)
专题命中 偏好对齐 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG
机构 * Gaoling School of Artificial Intelligence, Renmin University of China, Beijing(中国人民大学北京校区人工智能学院) ; Engineering Research Center of Next-Generation Intelligent Search(下一代智能搜索与推荐工程研究中心) ; Beijing Key Laboratory of Research on Large Models(北京大型模型研究重点实验室)
专题命中 偏好对齐 :alignment(abstract);分类 cs.AI
Comments Accepted By NeurIPS 2025
专题命中 安全训练 :safety(title)
Comments This work has been submitted to a conference for possible publication and is under review. Paper summary: 8 pages, 5 figures, 2 tables
机构 * Beijing University of Posts and Telecommunications(北京邮电大学) ; National University of Singapore(新加坡国立大学) ; CSIRO’s Data61(CSIRO数据61) ; Responsible AI Research (RAIR) Centre, The University of Adelaide(负责任人工智能研究(RAIR)中心,阿德莱德大学) ; Tsinghua University(清华大学)
专题命中 越狱攻击 :safety(title,abstract);jailbreak(abstract);分类 cs.LG
Comments Accepted to NeurIPS 2025
机构 * University of Cagliari(卡利亚里大学) ; Centre for AI Governance(人工智能治理中心) ; Foundation AI – Cisco Systems Inc.(AI基金会——思科系统公司)
专题命中 越狱攻击 :safety(abstract);jailbreak(abstract);分类 cs.CL、cs.AI、cs.LG
机构 * ELLIS Institute Tübingen(图宾根ELLIS研究所) ; MPI for Intelligent Systems Tübingen(图宾根智能系统研究所) ; AI Center(人工智能中心)
专题命中 提示注入 :prompt injection(title,abstract);分类 cs.LG
机构 * Department of Computer Science(计算机科学系) ; Virginia Tech(弗吉尼亚理工学院)
专题命中 幻觉与事实性 :alignment(abstract);分类 cs.LG
专题命中 安全评测 :alignment(title,abstract);分类 cs.CL、cs.AI
机构 * CDTI
专题命中 安全评测 :trustworthy(title);分类 cs.AI、cs.LG
Comments 16 pages, 4 figures
Journal ref IEEE Transactions on Neural Networks and Learning Systems, Early Access, 2025
机构 * NII LLMC(日本信息处理学会大语言模型中心) ; The University of Tokyo(东京大学) ; NAIST(日本科学技术大学) ; Nagoya Institute of Technology(名古屋技术大学)
专题命中 安全评测 :alignment(title);分类 cs.CL
Comments Accepted to EMNLP 2025 (Main Conference). Models and evaluation results available at: https://github.com/llm-jp/massive-sft
机构 * Oak Ridge National Laboratory(橡树岭国家实验室)
专题命中 安全评测 :safety(abstract);trustworthy(abstract);分类 cs.AI
Comments Preprint Submitted to ACM Transactions on AI for Science (TAIS)
机构 * Lane Department of Computer Science and Electrical Engineering, West Virginia University(计算机科学与电气工程系,西弗吉尼亚大学)
专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG
Journal ref Sustainable Energy, Grids and Networks, Vol. 44, December 2025, 102022
机构 * Indian Institute of Technology(印度理工学院)
专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.LG
机构 * Microsoft Responsible AI Research(微软负责任人工智能研究) ; University of California(加州大学)
专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.AI
机构 * University of Waterloo(滑铁卢大学) ; University of Oxford(牛津大学) ; Vector Institute(向量研究所)
专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI
Comments Project Page: https://github.com/Paper2Poster/Paper2Poster
机构 * KAIST(韩国科学技术院)
专题命中 安全评测 :alignment(abstract);分类 cs.AI
Comments 39th Conference on Neural Information Processing Systems (NeurIPS 2025) Workshop: The First Workshop on Generative and Protective AI for Content Creation
机构 * AI Singapore ; National University of Singapore(新加坡国立大学) ; Nanyang Technological University(南洋理工大学)
专题命中 安全评测 :alignment(abstract);分类 cs.CL
Comments Accepted at IJCNLP-AACL 2025 (Main Track). We released our model at https://huggingface.co/collections/aisingapore/sea-lionv3-672589a39cdadd6a5b199581
机构 * Case Western Reserve University(凯斯西储大学)
专题命中 安全评测 :trustworthy(abstract);分类 cs.LG
Comments Datasets link: https://huggingface.co/datasets/LLDDSS/Causal3D_Dataset
机构 * Indian Institute of Information Technology Dharwad, India(印度达拉瓦德信息科技学院) ; Indian Institute of Technology Indore, India(印度印度理工学院) ; Malaviya National Institute of Technology Jaipur, India(马拉维亚国家理工学院)
专题命中 安全评测 :alignment(abstract)
Comments Accepted at IEEE International Conference on Data Mining (ICDM) 2025
专题命中 安全评测 :alignment(abstract)
专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI、cs.CY
机构 * Seoul National University(首尔国立大学) ; Osaka University(大阪大学) ; Yonsei University(延世大学)
专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI
Comments Under review at The Web Conference 2026 (Semantics & Knowledge track). Code will be released upon acceptance. This arXiv v1 contains no repository links to preserve double-blind review
专题命中 其他安全 :alignment(title,abstract)
Comments Preprint under review at IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI), 2025
机构 * Tencent Youtu Lab(腾讯优图实验室)
专题命中 其他安全 :alignment(title,abstract)
Comments 12 pages, 7 figures
机构 * NOVA LINCS, NOVA School of Science and Technology(NOVA LINCS,NOVA科学与技术学院)
专题命中 其他安全 :alignment(title,abstract)
Comments Updated version: corresponds to the ACM MM '25 published paper and includes full appendix material
机构 * Humains AI Research(Humains人工智能研究) ; Inpris Ltd(Inpris公司)
专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG
机构 * Department of Linguistics University of Washington(语言学系华盛顿大学) ; Department of Linguistics Stanford University(语言学系斯坦福大学)
专题命中 其他安全 :alignment(abstract);分类 cs.CL
Comments 25 pages, 5 figures | EMNLP 2025 camera-ready version
机构 * School of Artificial Intelligence, Beijing University of Posts(人工智能学院,北京邮电大学) ; Beijing Big Data Center(北京大数据中心)
专题命中 其他安全 :safety(abstract);分类 cs.AI