SPINAL -- Scaling-law and Preference Integration in Neural Alignment Layers
SPINAL -- 神经对齐层中的缩放律与偏好整合
专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);DPO(abstract);分类 cs.CL、cs.AI、cs.LG
AI总结 SPINAL通过分析神经对齐层的几何特性,量化对齐在深度层的集中表现及稳定性变化。
AI 大模型
大模型对齐、安全、越狱、红队、提示注入和可信评测。
SPINAL -- 神经对齐层中的缩放律与偏好整合
专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);DPO(abstract);分类 cs.CL、cs.AI、cs.LG
AI总结 SPINAL通过分析神经对齐层的几何特性,量化对齐在深度层的集中表现及稳定性变化。
机构 * Department of Applied Mathematics and Statistics(应用数学与统计学系) ; Stony Brook University(石溪大学) ; Institute for Advanced Computational Science(先进计算科学研究所) ; Center of Excellence in Wireless and Information Technology (CEWIT)(无线与信息技术卓越中心) ; AI Innovation Institute(人工智能创新研究院) ; Bloomberg(彭博) ; Toronto, ON M5J 2S1(多伦多,ON M5J 2S1) ; San Francisco, CA 94105(旧金山,CA 94105) ; Bloomberg New York, NY 10022(纽约,NY 10022)
专题命中 偏好对齐 :DPO(title,abstract);alignment(abstract);RLHF(abstract);分类 cs.CL、cs.AI、cs.LG
专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);DPO(abstract);分类 cs.CL、cs.AI、cs.LG
机构 * Key Laboratory of Big Data & Artificial Intelligence in Transportation, (Beijing Jiaotong University), Ministry of Education(大数据与人工智能交通运输联合实验室,(北京交通大学)教育部) ; School of Computer Science and Technology, Beijing Jiaotong University, Beijing, China(计算机科学与技术学院,北京交通大学,北京,中国) ; Tencent Inc, China(腾讯公司,中国)
专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);DPO(abstract);分类 cs.CL、cs.AI、cs.LG
Comments ACL 2025 Main Conference, code available at: https://github.com/songmzhang/AlignDistil
机构 * Princeton University(普林斯顿大学)
专题命中 偏好对齐 :RLHF(title,abstract);alignment(abstract);DPO(abstract);分类 cs.CL、cs.AI、cs.LG
Comments 27 pages, 18 figures
机构 * ShanghaiTech University(上海科技大学) ; Henan University(河南大学) ; Liaoning University of Traditional Chinese Medicine(辽宁中医药大学)
专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);DPO(abstract);分类 cs.CL、cs.AI、cs.LG
专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);DPO(abstract);分类 cs.CL、cs.AI、cs.LG
机构 * University College London(伦敦大学学院) ; University of Bologna(博洛尼亚大学)
专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);DPO(abstract);分类 cs.AI、cs.CY、cs.LG
Comments Published at the 13th International Conference on Learning Representations (ICLR'25), Singapore, Apr 2025. https://openreview.net/forum?id=MeGDmZjUXy
机构 * The University of Texas at Austin(德克萨斯大学奥斯汀分校) ; Princeton Language & Intelligence(普林斯顿语言与智能)
专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);safety(abstract);分类 cs.CL、cs.AI、cs.LG
专题命中 偏好对齐 :safety(title,abstract);RLHF(abstract);harmlessness(abstract);分类 cs.CL、cs.AI、cs.LG
Comments The paper has been accepted in ICLR 2025 as spotlight presentation
专题命中 偏好对齐 :safety(title,abstract);alignment(abstract);harmlessness(abstract);分类 cs.AI、cs.CY、cs.LG
专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);DPO(abstract);分类 cs.CL、cs.AI、cs.LG
专题命中 偏好对齐 :DPO(title,abstract);alignment(abstract);safety(abstract);分类 cs.CL、cs.AI、cs.LG
专题命中 偏好对齐 :RLHF(title,abstract);alignment(abstract);DPO(abstract);分类 cs.CL、cs.AI、cs.LG
专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);DPO(abstract);分类 cs.CL、cs.AI、cs.LG
专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);DPO(abstract);分类 cs.CL、cs.AI、cs.LG
Comments Preprint Version
专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);DPO(abstract);分类 cs.CL、cs.AI、cs.LG
Comments 44 pages, 10 tables
专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);DPO(abstract);分类 cs.CL、cs.AI、cs.LG
Comments 27 pages, 4 figures, 5 tables
专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);DPO(abstract);分类 cs.CL、cs.AI、cs.LG
Comments 16 pages, 4 figures, Accepted to COLM 2024
专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);DPO(abstract);分类 cs.CL、cs.AI、cs.LG
Comments 24 pages, 6 figures, 3 tables
专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);DPO(abstract);分类 cs.CL、cs.CY、cs.LG
Comments Accepted to ACL 2024 main conference
专题命中 偏好对齐 :RLHF(title,abstract);alignment(abstract);DPO(abstract);分类 cs.CL、cs.AI、cs.LG
专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);DPO(abstract);分类 cs.CL、cs.AI、cs.LG
专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);DPO(abstract);分类 cs.AI、cs.LG
Comments accepted for AAAI 2026 Special Track on AI Alignment
机构 * University of Illinois Urbana–Champaign(伊利诺伊大学厄巴纳-香槟分校)
专题命中 偏好对齐 :safety(title,abstract);alignment(abstract);jailbreak(abstract);分类 cs.AI、cs.LG
Comments Previous title: "Saffron-1: Towards an Inference Scaling Paradigm for LLM Safety Assurance"
专题命中 偏好对齐 :alignment(title,abstract);DPO(abstract);safety(abstract);分类 cs.CL、cs.LG
Comments Pre-print. Submitted to the ICLR 2024 Workshop on Representational Alignment (Re-Align)
能思考、对话更出色的语言模型
机构 * Princeton Language and Intelligence(普林斯顿语言与智能研究所) ; Princeton University(普林斯顿大学)
专题命中 偏好对齐 :RLHF(summary_cn,abstract);DPO(abstract,abstract_cn);分类 cs.CL
AI总结 本文提出RLMT方法,将RLVR扩展至开放式任务,使语言模型生成长CoT推理,在多基准测试中优于RLHF,仅用少量提示训练的基础模型性能超多阶段后训练的指令微调模型
Comments COLM 2026; we release our code, data, and artifacts publicly at https://github.com/princeton-pli/RLMT
多语言情感感知文本摘要:一种用于一致性维护的强化学习方法
机构 * Instituto Politécnico Nacional (IPN), Centro de Investigación en Computación (CIC)(国立理工学院(IPN),计算研究中心(CIC))
专题命中 偏好对齐 :RLHF(summary_cn,abstract);alignment(abstract);safety(abstract);分类 cs.CL
AI总结 研究RLHF摘要中的情感漂移现象,提出基于策略归因框架的情感感知KL正则化方法,在保持摘要质量的同时缓解情感中性化。
为大型语言模型进行后训练的强化学习:综述
机构 * Salesforce ; AWS AI Labs ; Airbnb
专题命中 偏好对齐 :RLHF(summary_cn,abstract);DPO(abstract,abstract_cn);分类 cs.CL
AI总结 本文综述了通过强化学习进行大型语言模型后训练的方法,分析了RLHF和RLVR等技术,并提出统一的策略梯度框架,旨在为研究人员提供技术参考。
以自身声音对齐:用于减轻大型视觉-语言模型幻觉的自我纠正偏好学习
机构 * Graduate School of Advanced Imaging Sciences, Multimedia and Film(高级影像科学研究生院,多媒体与电影系) ; Department of Artificial Intelligence(人工智能系)
专题命中 偏好对齐 :DPO(summary_cn,abstract);alignment(abstract,abstract_cn);分类 cs.AI
AI总结 本文提出AVES-DPO框架,通过内在知识生成分布数据,利用共识验证机制诊断幻觉并引导模型自我纠正,有效缓解LVLMs的幻觉问题,仅需5200样本即优于现有基线。
Comments Accepted to ACL 2026