Towards Analyzing and Understanding the Limitations of DPO: A Theoretical Perspective
专题命中 偏好对齐 :DPO(title,abstract);分类 cs.CL、cs.AI
Comments Draft version
AI 大模型
大模型对齐、安全、越狱、红队、提示注入和可信评测。
专题命中 偏好对齐 :DPO(title,abstract);分类 cs.CL、cs.AI
Comments Draft version
专题命中 偏好对齐 :RLHF(title,abstract);分类 cs.CL、cs.AI
Comments Accepted to NAACL 2024. (7 pages)
专题命中 偏好对齐 :alignment(title,abstract);分类 cs.CL、cs.AI
Comments Accepted by LREC-COLING 2024
专题命中 偏好对齐 :alignment(title);RLHF(abstract);分类 cs.AI、cs.LG
专题命中 偏好对齐 :alignment(title,abstract);分类 cs.CL、cs.AI
Comments To appear in NeurIPS 2023 Workshop SyntheticData4ML
专题命中 偏好对齐 :alignment(title,abstract);分类 cs.CL、cs.AI
专题命中 偏好对齐 :RLHF(title,abstract);分类 cs.AI、cs.LG
专题命中 偏好对齐 :alignment(title,abstract);分类 cs.AI、cs.LG
专题命中 偏好对齐 :RLHF(title,abstract);分类 cs.CL、cs.AI
Comments 9 pages, working in progress
专题命中 偏好对齐 :alignment(title,abstract);分类 cs.AI、cs.LG
扩散模型对齐:基础、挑战与未来
机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) ; Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) ; Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学) ; Baidu Inc.(百度公司) ; The University of Bologna(博洛尼亚大学)
专题命中 偏好对齐 :alignment(title,abstract);分类 cs.LG
AI总结 本文综述了扩散模型对齐的基础、挑战及未来方向,探讨了对齐技术、评估方法及当前挑战的解决方案。
Comments Accepted at ACM Computing Surveys. 35 pages, 5 figures, 4 tables. Paper List: github.com/xie-lab-ml/awesome-alignment-of-diffusion-models
专题命中 偏好对齐 :DPO(title,abstract);分类 cs.CL
Comments 17 pages, 1 figures, 2 tables. Technical report. Introduces PureTC-1B, an adapter-based pipeline for stabilizing Small Language Models in Traditional Chinese using CPT, SFT, and DPO
机构 * Virginia Tech(弗吉尼亚理工大学) ; Amazon(亚马逊公司)
专题命中 偏好对齐 :alignment(title,abstract);分类 cs.LG
Comments Data-Centric Human Preference with Rationales for Direct Preference Alignment
机构 * UC Berkeley / Meta(伯克利大学/Meta) ; Meta
专题命中 偏好对齐 :prompt injection(title,abstract);分类 cs.LG
Comments ACM CCS 2025. Key words: prompt injection defense, LLM security, LLM-integrated applications
专题命中 偏好对齐 :alignment(title,abstract);分类 cs.CL
Comments Accepted at ICML 2024. This camera-ready version adds MT-Bench evaluations, a human study, more thorough analysis of length bias. Code at https://github.com/tml-epfl/long-is-more-for-alignment
BOUND:在搜索-控制边界处的摘要引导式校正偏好蒸馏
专题命中 偏好对齐 :DPO(summary_cn,abstract)
AI总结 针对大语言模型深度搜索智能体的持续漂移问题,提出BOUND框架,通过摘要引导构建偏好对,用DPO蒸馏偏好,在7个基准上多数数据集和指标优于基线。
Comments 15 pages
偏好优化导致LLM预测市场中的单一文化
专题命中 偏好对齐 :DPO(summary_cn,abstract)
AI总结 研究发现,使用直接偏好优化(DPO)微调的LLM代理在预测市场中产生高度相关的错误,导致集体预测能力远低于独立代理数量,且该现象由偏好优化驱动。
当距离干扰:BT损失中表示距离偏差对奖励模型的影响
机构 * University of California, Berkeley(加州大学伯克利分校)
专题命中 偏好对齐 :RLHF(abstract,abstract_cn);alignment(abstract);分类 cs.CL、cs.AI、cs.LG
AI总结 分析BT损失中表示距离导致的梯度偏差,提出NormBT自适应归一化方案,提升奖励模型在细粒度区分上的性能。
Comments ICML 2026
LongCat-Video-Avatar 1.5 技术报告
机构 * Meituan LongCat Team(美团LongCat团队)
专题命中 偏好对齐 :RLHF(summary_cn,abstract)
AI总结 本文提出 LongCat-Video-Avatar 1.5,一个通过升级音频编码器、优化训练策略、数据筛选和RLHF训练实现高精度唇同步、全身时间稳定性和长视频生成的开放框架,在多个基准测试中达到或超越商业系统性能。
Comments Homepage: https://meigen-ai.github.io/LongCat-Video-Avatar-1.5-Page/ Github: https://github.com/meituan-longcat/LongCat-Video
是时候正确了:提升视觉语言模型中的模拟时钟读取和指针空间推理能力
机构 * Incheon National University(延世国立大学) ; McGill University(麦吉尔大学)
专题命中 偏好对齐 :DPO(summary_cn,abstract)
AI总结 针对视觉语言模型在真实环境中读取模拟时钟的挑战,提出TickTockVQA数据集和Swap-DPO微调框架,显著提升时钟读取准确性和鲁棒性。
Comments Accepted to CVPR 2026 Findings
RLBFF:二进制灵活反馈用于连接人类反馈与可验证奖励
机构 * NVIDIA
专题命中 偏好对齐 :RLHF(abstract,abstract_cn);alignment(abstract);分类 cs.CL、cs.AI、cs.LG
AI总结 RLBFF结合人类偏好与规则验证,提升奖励模型对响应质量的精准捕捉,优于Bradley-Terry模型,在RM-Bench和JudgeBench上取得优异成绩,且支持用户自定义反馈原则。
Comments Published at ICLR 2026, 21 pages
硅哲学家中的异质性崩溃
机构 * Adobe Inc.(Adobe公司) ; Stanford University(斯坦福大学)
专题命中 偏好对齐 :DPO(abstract,abstract_cn);alignment(abstract);分类 cs.CL、cs.CY、cs.LG
AI总结 研究发现硅样本在哲学领域系统性地压缩了异质性,语言模型过度相关哲学判断,产生人工共识,影响对齐、评估和硅样本替代人类判断的使用。
大语言模型判官的保质期:未来证明、向后兼容性和问题泛化
机构 * Salesforce AI Research(Salesforce AI研究院) ; University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) ; Dartmouth College(达特茅斯学院)
专题命中 偏好对齐 :DPO(abstract,abstract_cn);alignment(abstract);分类 cs.CL、cs.AI、cs.LG
AI总结 本文研究了细调判官模型在现实部署中的挑战,探讨了未来证明、向后兼容性和问题泛化三个核心问题,并通过实验发现持续学习在适应响应分布变化方面表现更优。
Comments Updated after ICLR 2026 Acceptance; 29 pages;
Skywork-Reward-V2: 通过人机协同扩大偏好数据整理
机构 * Research, Skywork AI(2050研究, Skywork AI)
专题命中 偏好对齐 :alignment(abstract);RLHF(abstract);safety(abstract);分类 cs.CL、cs.AI、cs.LG
AI总结 Skywork-Reward-V2通过人机协同扩大偏好数据整理,提出SynPref-40M数据集和两阶段流程,训练出八种参数不同的奖励模型,在多个基准中取得最佳性能。
Comments ICLR 2026 Poster
RAPO:面向通用安全推理的风险感知优化
机构 * Shanghai AI Laboratory(上海人工智能实验室)
专题命中 偏好对齐 :alignment(abstract);safety(abstract);jailbreak(abstract);分类 cs.CL、cs.AI、cs.LG
AI总结 RAPO 提出了一种风险感知优化框架,通过提升安全推理的泛化能力,增强大型推理模型在复杂攻击下的安全性与实用性。
超越工具性和替代性范式:引入机器文化作为大型语言模型中的涌现现象
专题命中 偏好对齐 :alignment(abstract);RLHF(abstract);safety(abstract);分类 cs.CL、cs.AI、cs.CY
AI总结 本研究提出机器文化作为大型语言模型中的一种新兴现象,挑战传统工具性和替代性范式,揭示模型在文化表现上的独特特性。
Comments 16 pages, 6 figures
SecureCAI: 面向网络安全操作的抗注入大语言模型助手
机构 * Computer Science Department, Cybersecurity and Artificial Intelligence Division(计算机科学系、网络安全与人工智能 division)
专题命中 偏好对齐 :safety(abstract);prompt injection(abstract);trustworthy(abstract);AI safety(abstract)
AI总结 SecureCAI通过引入安全意识护栏和偏好优化,有效提升网络安全操作中对抗攻击的防御能力。
机构 * Stanford University(斯坦福大学)
专题命中 偏好对齐 :alignment(abstract);DPO(abstract);safety(abstract);分类 cs.CL、cs.AI、cs.LG
Comments COLM 2025 camera-ready version; 9 pages main text, 39 pages total
专题命中 偏好对齐 :alignment(abstract);RLHF(abstract);DPO(abstract);分类 cs.CL、cs.AI、cs.LG
专题命中 偏好对齐 :alignment(abstract);RLHF(abstract);DPO(abstract);分类 cs.CL、cs.AI、cs.LG
Comments 28 pages, 1 figure, 4 table