A Principled Loss Function for Direct Language Model Alignment
专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);DPO(abstract);分类 cs.AI、cs.LG
AI 大模型
大模型对齐、安全、越狱、红队、提示注入和可信评测。
专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);DPO(abstract);分类 cs.AI、cs.LG
机构 * University of Science and Technology of China(中国科学技术大学) ; Mach Drive(马车驱动)
专题命中 偏好对齐 :alignment(title,abstract);DPO(abstract)
机构 * Arizona State University(亚利桑那州立大学)
专题命中 偏好对齐 :alignment(title,abstract)
Comments EMNLP 2025 (Main)
机构 * LLM Department, Tencent(腾讯大模型部门) ; HunYuan Infra Team(文心一言基础设施团队) ; The Chinese University of Hong Kong(香港中文大学)
专题命中 偏好对齐 :RLHF(abstract);分类 cs.CL、cs.AI、cs.LG
Comments Work in progress
机构 * Preferred Networks Inc(Preferred Networks公司)
专题命中 偏好对齐 :DPO(abstract);分类 cs.CL、cs.AI、cs.LG
机构 * Peking University(北京大学) ; Tsinghua University(清华大学) ; Mila - Québec AI Institute(魁北克人工智能研究所)
专题命中 偏好对齐 :alignment(abstract);DPO(abstract)
Comments NeurIPS 2025. Code: https://github.com/Gen-Verse/HermesFlow
专题命中 偏好对齐 :DPO(abstract);分类 cs.CL、cs.AI
Comments 31 pages, 16 figures, 12 tables
机构 * Indian Institute of Technology Patna(印度帕纳杰大学) ; CRISIL LTD(CRISIL公司)
专题命中 偏好对齐 :DPO(abstract);分类 cs.AI
专题命中 安全训练 :safety(title,abstract);分类 cs.AI
Comments 14 pages, 10 figures
机构 * Harvard Medical School(哈佛医学院) ; Harvard College(哈佛学院) ; AstraZeneca(阿斯利康) ; Carnegie Mellon University(卡内基梅隆大学) ; Dana-Farber Cancer Institute and Harvard Medical School(达纳-法伯癌症研究所和哈佛医学院) ; Harvard University(哈佛大学) ; Broad Institute of MIT and Harvard(MIT和哈佛大学 Broad研究所) ; Harvard Data Science Initiative(哈佛数据科学计划)
专题命中 安全训练 :alignment(abstract);safety(abstract);分类 cs.AI、cs.LG
机构 * Indian Institute of Technology, Roorkee(印度理工学院拉胡尔分校)
专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG;trustworthy(comments)
Comments Accepted at ICCV2025 Workshop on Safe and Trustworthy Multimodal AI Systems
机构 * School of Computing and Augmented Intelligence, Arizona State University(计算与增强智能学院,亚利桑那州立大学) ; School of Mathematical and Statistical Sciences, Arizona State University(数学与统计科学学院,亚利桑那州立大学)
专题命中 安全训练 :safety(abstract);分类 cs.AI
专题命中 安全训练 :safety(abstract);分类 cs.LG
机构 * UiT The Arctic University of Norway(乌塔大学极地大学) ; Technical University of Denmark(技术大学) ; Østfold Hospital Trust(奥斯fold医院信托) ; Vestre Viken Hospital Trust(维斯特维肯医院信托) ; Radboud University Nijmegen Medical Centre(拉德堡德大学奈梅亨医疗中心) ; The Netherlands Cancer Institute(荷兰癌症研究所) ; Antoni van Leeuwenhoek Hospital(安东尼·弗莱明医院) ; University of Copenhagen(哥本哈根大学)
专题命中 安全训练 :alignment(abstract)
机构 * Fudan University(复旦大学) ; Huawei Technologies Ltd.(华为技术有限公司) ; Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
专题命中 越狱攻击 :alignment(abstract);safety(abstract);jailbreak(abstract);分类 cs.AI
机构 * College of Computer Science and Technology, Zhejiang University(浙江大学计算机科学与技术学院) ; Palo Alto Networks(帕洛阿尔托网络公司) ; School of Software Technology, Zhejiang University(浙江大学软件技术学院)
专题命中 红队测试 :red teaming(title,abstract);分类 cs.AI
机构 * Institute for Machine Learning and Analytics (IMLA)(机器学习与分析研究所)
专题命中 提示注入 :prompt injection(title,abstract);分类 cs.LG
机构 * Independent Researcher(独立研究者)
专题命中 安全评测 :safety(title);分类 cs.AI、cs.CY、cs.LG
Comments 9 pages plus citations and appendix, 7 figures
机构 * State Key Laboratory of Virtual Reality Technology and Systems, School of Computer Science and Engineering, Beihang University, Beijing, China(虚拟现实技术与系统国家重点实验室,计算机科学与工程学院,北京航空航天大学) ; Hangzhou Innovation Institute of Beihang University, Zhejiang Key Laboratory of Industrial Big Data and Robot Intelligent Systems, Hangzhou, China(北京航空航天大学杭州创新研究院,浙江省工业大数据与机器人智能系统重点实验室) ; Center for AI Business Innovation, Department of Management Science and Systems, University at Buffalo, Buffalo, New York, USA(人工智能商业创新中心,管理科学与系统系,布法罗大学) ; University of North Texas, Denton, Texas, USA(德克萨斯大学达文波特分校) ; Zhongguancun Laboratory, Beijing, China(中关村实验室)
专题命中 安全评测 :alignment(abstract);safety(abstract);分类 cs.CL、cs.LG
Comments 28 pages, 32 figures, accepted to the Findings of EMNLP 2025
机构 * Synkrasis Labs(Synkrasis实验室) ; Harbin Institute of Technology(哈尔滨工业大学)
专题命中 安全评测 :alignment(abstract);safety(abstract);分类 cs.AI
Comments Accepted: LAW 2025 Workshop NeurIPS 2025
机构 * University College London(伦敦大学学院)
专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI
专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG
Comments 18 pages, 21 figures
机构 * Harbin Institute of Technology(哈尔滨工业大学)
专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.AI
机构 * DISI, University of Trento(特伦托大学DISI中心) ; MaiNLP, Center for Information and Language Processing, LMU Munich(慕尼黑大学信息与语言处理中心) ; Munich Center for Machine Learning (MCML), Munich, Germany(慕尼黑机器学习中心) ; Free University of Bozen-Bolzano, Italy(博兹纳自由大学)
专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI
Comments EMNLP 2025 Main, 38 pages, 33 figures
机构 * Stability AI ; SketchX, University of Surrey(SketchX,大学)
专题命中 安全评测 :alignment(abstract);分类 cs.AI
Comments Project Page: https://hmrishavbandy.github.io/sd35flash/
机构 * University of California, Berkeley(加州大学伯克利分校)
专题命中 安全评测 :alignment(abstract);分类 cs.AI
Comments 39th Conference on Neural Information Processing Systems (NeurIPS 2025) Workshop: Evaluating the Evolving LLM Lifecycle: Benchmarks, Emergent Abilities, and Scaling
机构 * Peking University(北京大学) ; LLM-Core
专题命中 安全评测 :trustworthy(abstract);分类 cs.CL
专题命中 安全评测 :trustworthy(abstract);分类 cs.AI
Comments 11 pages, 2 figures, 2 tables
机构 * Embedded Systems Group (Dept.-E)(嵌入式系统组) ; Virtual Vehicle Research GmbH(虚拟车辆研究公司) ; Control Systems Group (Dept.-E)(控制系统组) ; Institute of Visual Computing(视觉计算研究所) ; Graz University of Technology(格拉茨技术大学)
专题命中 安全评测 :safety(abstract)
Comments Accepted for presentation at the 2025 IEEE International Automated Vehicle Validation Conference (IAVVC 2025). Final version to appear in IEEE Xplore
专题命中 AI治理与伦理 :safety(abstract);AI safety(abstract);分类 cs.CL、cs.AI、cs.CY