InfAlign: Inference-aware language model alignment
机构 * Google DeepMind(谷歌DeepMind) ; Google Research(谷歌研究)
专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);分类 cs.CL、cs.LG
AI 大模型
大模型对齐、安全、越狱、红队、提示注入和可信评测。
机构 * Google DeepMind(谷歌DeepMind) ; Google Research(谷歌研究)
专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);分类 cs.CL、cs.LG
专题命中 偏好对齐 :alignment(title,abstract)
机构 * Peking University(北京大学) ; Tsinghua University(清华大学)
专题命中 偏好对齐 :safety(abstract);分类 cs.AI、cs.LG
Comments 10 pages, 9 figures, Under review as a full paper at AAAI 2026. A preliminary version is under review at the NeurIPS 2025 Workshop on Reliable ML from Unreliable Data
专题命中 偏好对齐 :DPO(abstract);分类 cs.AI
机构 * University of Ljubljana, Faculty of Computer and Information Science(卢布尔雅那大学计算机与信息科学系)
专题命中 偏好对齐 :DPO(abstract);分类 cs.CL
Comments Paper with individual presentation at LUHME workshop at ECAI 2025
专题命中 偏好对齐 :alignment(abstract)
Comments Accepted by The 38th Annual ACM Symposium on User Interface Software and Technology (UIST Adjunct '25), September 28-October 1, 2025, Busan, Republic of Korea
机构 * Shuang Ao ; Gopal Rumchurn
专题命中 安全训练 :alignment(abstract);safety(abstract);分类 cs.AI
Comments 9 pages, 2 figures
机构 * Department of Electrical and Computer Engineering, Queen’s University(电气与计算机工程系,皇后大学) ; School of Computing, Queen’s University(计算机学院,皇后大学)
专题命中 越狱攻击 :jailbreak(title,abstract);alignment(abstract);safety(abstract);分类 cs.LG
机构 * University of Waterloo(多伦多大学) ; University of California, Berkeley(加州大学伯克利分校) ; University of Toronto(多伦多大学)
专题命中 越狱攻击 :alignment(abstract);分类 cs.CL
专题命中 越狱攻击 :prompt injection(abstract)
机构 * MIT CSAIL(麻省理工学院计算机科学与人工智能实验室) ; Sony(索尼) ; Weizmann Institute of Science(魏茨曼科学研究院) ; JKU(约翰纳斯堡大学) ; Tübingen AI Center(图宾根人工智能中心) ; UVA(乌得勒支大学) ; UNSW(新南威尔士大学) ; TU Graz(格拉茨技术大学) ; MIT-IBM(麻省理工学院-IBM)
专题命中 越狱攻击 :safety(abstract)
Comments Code: https://github.com/jmiemirza/GLOV
专题命中 红队测试 :red teaming(title,abstract)
专题命中 提示注入 :prompt injection(title,abstract);分类 cs.CL、cs.AI
Comments EMNLP 2025
机构 * College of Computer Science and Engineering, Northeastern University(计算机科学与工程学院,东北大学) ; Distributed Systems Group, TU Wien(分布式系统组,TU Wien) ; College of Software, Northeastern University(软件学院,东北大学) ; College of Information Science and Engineering, Northeastern University(信息科学与工程学院,东北大学)
专题命中 隐私与版权 :alignment(abstract);分类 cs.AI、cs.LG
专题命中 安全评测 :trustworthy(title,abstract);分类 cs.CL、cs.AI
机构 * Laboratory of Brain Atlas and Brain-inspired Intelligence, Institute of Automation Chinese Academy of Sciences (CASIA)(中国科学院自动化研究所脑图谱与类脑智能实验室) ; School of Artificial Intelligence, University of Chinese Academy of Sciences (UCAS)(中国科学院大学人工智能学院) ; School of Systems Science, Beijing Normal University(北京师范大学系统科学学院) ; School of Psychological and Cognitive Sciences & Beijing Key Laboratory of Behavior and Mental Health, Peking University(北京大学心理与认知科学学院) ; IDG/McGovern Institute for Brain Research, Peking University(北京大学IDG/ McGovern脑科学研究院) ; Institute for Artificial Intelligence & Key Laboratory of Machine Perception (Ministry of Education), Peking University(北京大学人工智能研究所) ; School of Future Technology, University of Chinese Academy of Sciences (UCAS)(中国科学院大学未来技术学院)
专题命中 安全评测 :alignment(title,abstract);分类 cs.CL、cs.AI
机构 * Engineering Department, School of Science and Technology (SST), City University of London(伦敦城市大学科学与技术学院工程系)
专题命中 安全评测 :safety(title,abstract);分类 cs.AI
机构 * Shanghai AI Lab(上海人工智能实验室) ; East China Normal University(东华大学)
专题命中 安全评测 :safety(title,abstract);分类 cs.CL
Comments Code and dataset are available at https://github.com/yangyangyang127/SafetyFlow
机构 * Department of Computer Science and Technology, College of AI, Institute for AI, Tsinghua-Bosch Joint ML Center, THBI Lab, BNRist Center, Tsinghua University(计算机科学与技术系、人工智能学院、人工智能研究所、清华-博世联合机器学习中心、THBI实验室、BNRist中心、清华大学) ; Institute of Artificial Intelligence, Beihang University(人工智能研究院、北航) ; RealAI
专题命中 安全评测 :alignment(abstract);safety(abstract);分类 cs.CL、cs.AI
Comments For Appendix, please refer to arXiv:2406.07057
专题命中 安全评测 :alignment(abstract);safety(abstract);分类 cs.LG
机构 * Ningbo Key Laboratory of Spatial Intelligence and Digital Derivative, Institute of Digital Twin, EIT(宁波空间智能与数字衍生关键实验室,数字孪生研究院,EIT) ; Logic Intelligence Technology(逻辑智能技术) ; BUPT(北京邮电大学) ; Xiamen University(厦门大学)
专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG
机构 * Department of Computing Science University of Alberta(计算科学系阿尔伯塔大学) ; ServiceNow Research(ServiceNow研究)
专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG
机构 * School of Electrical and Computer Engineering, University of Sydney(悉尼大学电气与计算机工程学院) ; School of Computer Science, University of Adelaide(阿德莱德大学计算机科学学院) ; School of Computing and Information Technology, University of Wollongong(沃林根大学计算与信息科技学院)
专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI
机构 * University of California at Berkeley(加州大学伯克利分校) ; Indian Institute of Technology Bombay(印度班加罗尔理工学院) ; Chalmers University of Technology and University of Gothenburg(查尔姆斯理工大学和哥德堡大学)
专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG
Comments This work has been accepted at ATVA'25
机构 * Otto von Guericke University Magdeburg(奥托·冯·格里克大学马格德堡)
专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.CY
专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.CY
Comments Accepted to AAAI/ACM AIES 2025
专题命中 安全评测 :alignment(abstract);分类 cs.LG
Comments Published at COLM 2025
机构 * Dept. of Computer Engineering(计算机工程系) ; Jamia Millia Islamia(Jamia Millia Islamia大学)
专题命中 安全评测 :trustworthy(abstract);分类 cs.AI
机构 * Department of Computer Science(计算机科学系) ; Czech Technical University Prague(捷克技术大学布拉格) ; Global Priorities Institute(全球优先研究所) ; University of Oxford(牛津大学) ; Foundations of Cooperative AI Lab(合作人工智能基础实验室) ; Carnegie Mellon University(卡内基梅隆大学)
专题命中 安全评测 :safety(abstract);分类 cs.AI
机构 * Qorvex Consulting ; Kleiner Perkins ; Wentworth Institute of Higher Education
专题命中 安全评测 :trustworthy(abstract);分类 cs.AI