Cost-Minimized Label-Flipping Poisoning Attack to LLM Alignment
专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);DPO(abstract);分类 cs.AI、cs.LG
Comments accepted for AAAI 2026 Special Track on AI Alignment
AI 大模型
大模型对齐、安全、越狱、红队、提示注入和可信评测。
专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);DPO(abstract);分类 cs.AI、cs.LG
Comments accepted for AAAI 2026 Special Track on AI Alignment
机构 * MIT CSAIL(麻省理工学院计算机科学与人工智能实验室)
专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);DPO(abstract);分类 cs.CL
Journal ref 13th International Conference on Learning Representations (ICLR 2025)
专题命中 偏好对齐 :alignment(title,abstract);DPO(abstract);safety(abstract)
机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) ; NewsBreak
专题命中 偏好对齐 :trustworthy(title);safety(abstract);分类 cs.CL
Comments Preprint update
专题命中 偏好对齐 :DPO(abstract);分类 cs.CL、cs.AI
Comments accepted to EMNLP2025
机构 * LinkedIn Corporation(LinkedIn公司) ; University of California San Diego(加州大学圣地亚哥分校)
专题命中 偏好对齐 :alignment(abstract);分类 cs.AI、cs.LG
专题命中 偏好对齐 :alignment(abstract);分类 cs.CL
专题命中 安全训练 :safety(title,abstract);AI safety(title,abstract);分类 cs.CY
专题命中 安全训练 :alignment(abstract)
机构 * Soongsil University(首尔大学)
专题命中 越狱攻击 :jailbreak(abstract);分类 cs.CL、cs.AI
专题命中 越狱攻击 :alignment(abstract);分类 cs.AI
机构 * Department of Computer Science, University of Chicago(芝加哥大学计算机科学系) ; Meta ; Department of Computer Science, University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校计算机科学系)
专题命中 红队测试 :red teaming(title,abstract);分类 cs.AI、cs.LG
专题命中 提示注入 :prompt injection(title,abstract);alignment(abstract);safety(abstract);分类 cs.AI
Comments 10 pages, 6 figures
专题命中 提示注入 :prompt injection(title,abstract);分类 cs.CL、cs.AI、cs.LG
Comments Published in USENIX Security Symposium 2024; the model sizes for closed-source models are from blog posts. For slides, see https://people.duke.edu/~zg70/code/PromptInjection.pdf
专题命中 提示注入 :prompt injection(title,abstract);分类 cs.AI
Comments Distinguished Paper Award in IEEE Symposium on Security and Privacy, 2025. For slides, see https://people.duke.edu/~zg70/code/PromptInjection.pdf
机构 * ETSIAE-UPM - School of Aeronautics, Universidad Politécnica de Madrid(西班牙马德里理工大学航空学院) ; Institute for Cross-Disciplinary Physics and Complex Systems (IFISC), CSIC-UIB(跨学科物理与复杂系统研究所(IFISC)) ; Center for Computational Simulation, Universidad Politécnica de Madrid, Campus de Montegancedo, Boadilla del Monte(马德里理工大学计算模拟中心)
专题命中 幻觉与事实性 :safety(abstract);分类 cs.LG
Comments 34 pages, 17 figures
Journal ref Engineering Failure Analysis, Volume 184, 2026, 110334
机构 * Tencent(腾讯)
专题命中 安全评测 :alignment(abstract);safety(abstract);分类 cs.CL、cs.AI
机构 * Dept. of Computer Science, The University of Texas at El Paso(得克萨斯大学埃尔帕索分校计算机科学系) ; Dept. of Mathematics, Mountain View High School(山景高中数学系) ; Dept. of Mathematics, El Paso Independent School District(埃尔帕索独立学区数学系)
专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.CY、cs.LG
机构 * Hong Kong Baptist University(香港 Baptist 大学) ; Beijing Normal-Hong Kong Baptist University(北京师范大学-香港 Baptist 大学) ; National University of Singapore(新加坡国立大学) ; Beijing Normal University(北京师范大学) ; Shanghai Jiao Tong University(上海交通大学)
专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.AI
Comments 28 pages, 14 figures, 19 tables
机构 * Microsoft Research(微软研究院) ; Northwestern University(西北大学)
专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.LG
机构 * Johnson & Johnson MedTech(强生医疗科技)
专题命中 安全评测 :alignment(abstract);分类 cs.AI
机构 * The Chinese University of Hong Kong(香港中文大学) ; Huawei Noah’s Ark Lab(华为诺亚实验室) ; The University of Hong Kong(香港大学)
专题命中 安全评测 :alignment(abstract);分类 cs.CL
Comments under review
机构 * Beijing Advanced Innovation Center for Future Blockchain and Privacy Computing(未来区块链与隐私计算先进创新中心) ; Institute of Artificial Intelligence, Beihang University(北京航空航天大学人工智能学院) ; Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所)
专题命中 安全评测 :safety(abstract);分类 cs.CL
Comments EMNLP2025Findings
机构 * Nanyang Technological University (NTU)(南洋理工大学) ; MiroMind(米罗Mind) ; Institute for Infocomm Research (I 2 R)(信息与通信研究院) ; A*STAR(科技研究局)
专题命中 安全评测 :alignment(abstract);分类 cs.CL
Comments Link: https://github.com/AudioLLMs/AudioBench/tree/main/IFEval-Audio
机构 * Centre for Responsible AI, IIT Madras(责任人工智能中心,IIT马德拉斯)
专题命中 安全评测 :trustworthy(abstract);分类 cs.CL
机构 * Bio LIMS INC(Bio LIMS公司)
专题命中 安全评测 :safety(abstract);分类 cs.AI
Comments 12 pages, 0 figures
专题命中 安全评测 :trustworthy(abstract);分类 cs.CL
专题命中 安全评测 :alignment(abstract)
机构 * School of Artificial Intelligence and Automation, Huazhong University of Science and Technology(人工智能与自动化学院,华中科技大学) ; Institute of Artificial Intelligence, Huazhong University of Science and Technology(人工智能研究院,华中科技大学) ; Hubei Key Laboratory of Brain-Inspired Intelligent Systems, Huazhong University of Science and Technology(湖北省脑启发智能系统重点实验室,华中科技大学) ; Key Laboratory of Image Processing and Intelligent Control (Huazhong University of Science and Technology), Ministry of Education(图像处理与智能控制重点实验室(华中科技大学),教育部) ; Key Laboratory of Intelligent Computing and Signal Processing, Ministry of Education, Anhui University(智能计算与信号处理重点实验室(安徽大学),教育部) ; College of Computing and Data Science (CCDS), Nanyang Technological University(计算与数据科学学院(CCDS),南洋理工大学) ; East China University of Science and Technology(东华大学)
专题命中 安全评测 :safety(abstract)
机构 * Leipzig University(莱比锡大学) ; Technical University Dresden(德累斯顿技术大学) ; University of Göttingen(哥廷根大学)
专题命中 AI治理与伦理 :trustworthy(abstract,comments);分类 cs.AI、cs.CY
Comments Accepted paper - ESORICS 2025 - International Workshop on Secure and Trustworthy Machine Unlearning Systems (STMUS)