arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 3266 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 偏好对齐 3266 篇

2404.04626 2024-04-09 cs.CL cs.AI 81%

Towards Analyzing and Understanding the Limitations of DPO: A Theoretical Perspective

Duanyu Feng, Bowen Qin, Chen Huang, Zheng Zhang, Wenqiang Lei

专题命中 偏好对齐 :DPO(title,abstract);分类 cs.CL、cs.AI

Comments Draft version

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.05553 2024-04-09 cs.CL cs.AI 81%

Removing RLHF Protections in GPT-4 via Fine-Tuning

Qiusi Zhan, Richard Fang, Rohan Bindu, Akul Gupta, Tatsunori Hashimoto, Daniel Kang

专题命中 偏好对齐 :RLHF(title,abstract);分类 cs.CL、cs.AI

Comments Accepted to NAACL 2024. (7 pages)

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.11124 2024-04-02 cs.CL cs.AI 81%

Scaling Data Diversity for Fine-Tuning Language Models in Human Alignment

Feifan Song, Bowen Yu, Hao Lang, Haiyang Yu, Fei Huang, Houfeng Wang, Yongbin Li

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.CL、cs.AI

Comments Accepted by LREC-COLING 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.11971 2023-12-27 cs.LG cs.AI 81%

Improving Generalization of Alignment with Human Preferences through Group Invariant Learning

Rui Zheng, Wei Shen, Yuan Hua, Wenbin Lai, Shihan Dou, Yuhao Zhou, Zhiheng Xi, Xiao Wang, Haoran Huang, Tao Gui, Qi Zhang, Xuanjing Huang

专题命中 偏好对齐 :alignment(title);RLHF(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.20033 2023-11-06 cs.CL cs.AI 81%

Synthetic Imitation Edit Feedback for Factual Alignment in Clinical Summarization

Prakamya Mishra, Zonghai Yao, Shuwei Chen, Beining Wang, Rohan Mittal, Hong Yu

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.CL、cs.AI

Comments To appear in NeurIPS 2023 Workshop SyntheticData4ML

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.06450 2023-10-12 cs.CL cs.AI 81%

Constructive Large Language Models Alignment with Diverse Feedback

Tianshu Yu, Ting-En Lin, Yuchuan Wu, Min Yang, Fei Huang, Yongbin Li

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.06147 2023-10-11 cs.LG cs.AI 81%

Reinforcement Learning in the Era of LLMs: What is Essential? What is needed? An RL Perspective on RLHF, Prompting, and Beyond

Hao Sun

专题命中 偏好对齐 :RLHF(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.05871 2023-10-10 cs.AI cs.LG cs.SY eess.SY 81%

Dynamic value alignment through preference aggregation of multiple objectives

Marcin Korecki, Damian Dailisan, Cesare Carissimo

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.10202 2023-09-20 cs.CL cs.AI 81%

Stabilizing RLHF through Advantage Model and Selective Rehearsal

Baolin Peng, Linfeng Song, Ye Tian, Lifeng Jin, Haitao Mi, Dong Yu

专题命中 偏好对齐 :RLHF(title,abstract);分类 cs.CL、cs.AI

Comments 9 pages, working in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
1811.07871 2018-11-20 cs.LG cs.AI cs.NE stat.ML 81%

Scalable agent alignment via reward modeling: a research direction

Jan Leike, David Krueger, Tom Everitt, Miljan Martic, Vishal Maini, Shane Legg

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.07253 2026-02-06 cs.LG cs.CV 80%

Alignment of Diffusion Models: Fundamentals, Challenges, and Future

扩散模型对齐:基础、挑战与未来

Buhua Liu, Shitong Shao, Bao Li, Lichen Bai, Zhiqiang Xu, Haoyi Xiong, James Kwok, Sumi Helal, Zeke Xie

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学) Baidu Inc.(百度公司) The University of Bologna(博洛尼亚大学)

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.LG

AI总结 本文综述了扩散模型对齐的基础、挑战及未来方向,探讨了对齐技术、评估方法及当前挑战的解决方案。

Comments Accepted at ACM Computing Surveys. 35 pages, 5 figures, 4 tables. Paper List: github.com/xie-lab-ml/awesome-alignment-of-diffusion-models

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01616 2025-10-03 cs.CL 80%

Efficient Training of Robust Traditional Chinese LLaMA-1B on a Single Consumer GPU: Continual Pre-training, SFT, and DPO

Yu-Cheng Chih, Ming-Tao Duan, Yong-Hao Hou

专题命中 偏好对齐 :DPO(title,abstract);分类 cs.CL

Comments 17 pages, 1 figures, 2 tables. Technical report. Introduces PureTC-1B, an adapter-based pipeline for stabilizing Small Language Models in Traditional Chinese using CPT, SFT, and DPO

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.14477 2025-07-15 cs.LG 80%

Data-Centric Human Preference with Rationales for Direct Preference Alignment

Hoang Anh Just, Ming Jin, Anit Sahu, Huy Phan, Ruoxi Jia

机构 * Virginia Tech(弗吉尼亚理工大学) Amazon(亚马逊公司)

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.LG

Comments Data-Centric Human Preference with Rationales for Direct Preference Alignment

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.05451 2025-07-04 cs.CR cs.LG 80%

SecAlign: Defending Against Prompt Injection with Preference Optimization

Sizhe Chen, Arman Zharmagambetov, Saeed Mahloujifar, Kamalika Chaudhuri, David Wagner, Chuan Guo

机构 * UC Berkeley / Meta(伯克利大学/Meta) Meta

专题命中 偏好对齐 :prompt injection(title,abstract);分类 cs.LG

Comments ACM CCS 2025. Key words: prompt injection defense, LLM security, LLM-integrated applications

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.04833 2024-06-05 cs.CL 80%

Long Is More for Alignment: A Simple but Tough-to-Beat Baseline for Instruction Fine-Tuning

Hao Zhao, Maksym Andriushchenko, Francesco Croce, Nicolas Flammarion

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.CL

Comments Accepted at ICML 2024. This camera-ready version adds MT-Bench evaluations, a human study, more thorough analysis of length bias. Code at https://github.com/tml-epfl/long-is-more-for-alignment

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.08768 2026-08-11 cs.IR 新提交 80%

BOUND: Brief-Guided Corrective Preference Distillation at Search-Control Boundaries

BOUND:在搜索-控制边界处的摘要引导式校正偏好蒸馏

Qingying Niu, Ruiyang Ren, Wayne Xin Zhao, Yaliang Li

专题命中 偏好对齐 :DPO(summary_cn,abstract)

AI总结 针对大语言模型深度搜索智能体的持续漂移问题,提出BOUND框架,通过摘要引导构建偏好对,用DPO蒸馏偏好,在7个基准上多数数据集和指标优于基线。

Comments 15 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.26583 2026-06-26 cs.CE 新提交 80%

Preference Optimization Drives Monoculture in LLM Prediction Markets

偏好优化导致LLM预测市场中的单一文化

James Begin, Brendan Gho, Suman Muppavarapu, Tyson Tsay, Atharva Mohan, Afnan Shaik, Ruizhe Li, Vasu Sharma, Archana Vaidheeswaran

专题命中 偏好对齐 :DPO(summary_cn,abstract)

AI总结 研究发现,使用直接偏好优化(DPO)微调的LLM代理在预测市场中产生高度相关的错误,导致集体预测能力远低于独立代理数量,且该现象由偏好优化驱动。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.06343 2026-06-10 cs.LG cs.AI cs.CL 版本更新 80%

When Distance Distracts: Representation Distance Bias in BT-Loss for Reward Models

当距离干扰:BT损失中表示距离偏差对奖励模型的影响

Tong Xie, Andrew Bai, Yuanhao Ban, Yunqi Hong, Haoyu Li, Cho-Jui Hsieh

机构 * University of California, Berkeley(加州大学伯克利分校)

专题命中 偏好对齐 :RLHF(abstract,abstract_cn);alignment(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 分析BT损失中表示距离导致的梯度偏差,提出NormBT自适应归一化方案,提升奖励模型在细粒度区分上的性能。

Comments ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.26486 2026-05-27 cs.CV 80%

LongCat-Video-Avatar 1.5 Technical Report

LongCat-Video-Avatar 1.5 技术报告

Meituan LongCat Team, Xunliang Cai, Meng Cheng, Feng Gao, Zhe Kong, Jiamu Li, Le Li, Weiheng Li, Hongyu Liu, Shuai Tan, Xiaoming Wei, Tianyu Yang, Yong Zhang

机构 * Meituan LongCat Team(美团LongCat团队)

专题命中 偏好对齐 :RLHF(summary_cn,abstract)

AI总结 本文提出 LongCat-Video-Avatar 1.5,一个通过升级音频编码器、优化训练策略、数据筛选和RLHF训练实现高精度唇同步、全身时间稳定性和长视频生成的开放框架,在多个基准测试中达到或超越商业系统性能。

Comments Homepage: https://meigen-ai.github.io/LongCat-Video-Avatar-1.5-Page/ Github: https://github.com/meituan-longcat/LongCat-Video

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.08011 2026-05-26 cs.CV 80%

It's Time to Get It Right: Improving Analog Clock Reading and Clock-Hand Spatial Reasoning in Vision-Language Models

是时候正确了:提升视觉语言模型中的模拟时钟读取和指针空间推理能力

Jaeha Choi, Jin Won Lee, Siwoo You, Jangho Lee

机构 * Incheon National University(延世国立大学) McGill University(麦吉尔大学)

专题命中 偏好对齐 :DPO(summary_cn,abstract)

AI总结 针对视觉语言模型在真实环境中读取模拟时钟的挑战,提出TickTockVQA数据集和Swap-DPO微调框架,显著提升时钟读取准确性和鲁棒性。

Comments Accepted to CVPR 2026 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21319 2026-05-19 cs.CL cs.AI cs.LG 80%

RLBFF: Binary Flexible Feedback to bridge between Human Feedback & Verifiable Rewards

RLBFF:二进制灵活反馈用于连接人类反馈与可验证奖励

Zhilin Wang, Jiaqi Zeng, Olivier Delalleau, Ellie Evans, Daniel Egert, Hoo-Chang Shin, Felipe Soares, Yi Dong, Oleksii Kuchaiev

机构 * NVIDIA

专题命中 偏好对齐 :RLHF(abstract,abstract_cn);alignment(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 RLBFF结合人类偏好与规则验证,提升奖励模型对响应质量的精准捕捉,优于Bradley-Terry模型,在RM-Bench和JudgeBench上取得优异成绩,且支持用户自定义反馈原则。

Comments Published at ICLR 2026, 21 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.23575 2026-04-30 cs.CY cs.CL cs.LG 80%

The Collapse of Heterogeneity in Silicon Philosophers

硅哲学家中的异质性崩溃

Yuanming Shi, Andreas Haupt

机构 * Adobe Inc.(Adobe公司) Stanford University(斯坦福大学)

专题命中 偏好对齐 :DPO(abstract,abstract_cn);alignment(abstract);分类 cs.CL、cs.CY、cs.LG

AI总结 研究发现硅样本在哲学领域系统性地压缩了异质性,语言模型过度相关哲学判断,产生人工共识,影响对齐、评估和硅样本替代人类判断的使用。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23542 2026-04-21 cs.CL cs.AI cs.LG 80%

On the Shelf Life of Fine-Tuned LLM-Judges: Future-Proofing, Backward-Compatibility, and Question Generalization

大语言模型判官的保质期:未来证明、向后兼容性和问题泛化

Janvijay Singh, Austin Xu, Yilun Zhou, Yefan Zhou, Dilek Hakkani-Tur, Shafiq Joty

机构 * Salesforce AI Research(Salesforce AI研究院) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Dartmouth College(达特茅斯学院)

专题命中 偏好对齐 :DPO(abstract,abstract_cn);alignment(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文研究了细调判官模型在现实部署中的挑战,探讨了未来证明、向后兼容性和问题泛化三个核心问题,并通过实验发现持续学习在适应响应分布变化方面表现更优。

Comments Updated after ICLR 2026 Acceptance; 29 pages;

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.01352 2026-03-04 cs.CL cs.AI cs.LG 80%

Skywork-Reward-V2: Scaling Preference Data Curation via Human-AI Synergy

Skywork-Reward-V2: 通过人机协同扩大偏好数据整理

Chris Yuhao Liu, Liang Zeng, Yuzhen Xiao, Jujie He, Jiacai Liu, Chaojie Wang, Rui Yan, Wei Shen, Fuxiang Zhang, Jiacheng Xu, Yang Liu, Yahui Zhou

机构 * Research, Skywork AI(2050研究, Skywork AI)

专题命中 偏好对齐 :alignment(abstract);RLHF(abstract);safety(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 Skywork-Reward-V2通过人机协同扩大偏好数据整理,提出SynPref-40M数据集和两阶段流程,训练出八种参数不同的奖励模型,在多个基准中取得最佳性能。

Comments ICLR 2026 Poster

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.04224 2026-02-05 cs.LG cs.AI cs.CL cs.CR math.OC 80%

RAPO: Risk-Aware Preference Optimization for Generalizable Safe Reasoning

RAPO:面向通用安全推理的风险感知优化

Zeming Wei, Qiaosheng Zhang, Xia Hu, Xingcheng Xu

机构 * Shanghai AI Laboratory(上海人工智能实验室)

专题命中 偏好对齐 :alignment(abstract);safety(abstract);jailbreak(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 RAPO 提出了一种风险感知优化框架,通过提升安全推理的泛化能力,增强大型推理模型在复杂攻击下的安全性与实用性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.17096 2026-01-27 cs.CY cs.AI cs.CL 80%

Beyond Instrumental and Substitutive Paradigms: Introducing Machine Culture as an Emergent Phenomenon in Large Language Models

超越工具性和替代性范式:引入机器文化作为大型语言模型中的涌现现象

Yueqing Hu, Xinyang Peng, Yukun Zhao, Lin Qiu, Ka-lai Hung, Kaiping Peng

专题命中 偏好对齐 :alignment(abstract);RLHF(abstract);safety(abstract);分类 cs.CL、cs.AI、cs.CY

AI总结 本研究提出机器文化作为大型语言模型中的一种新兴现象,挑战传统工具性和替代性范式,揭示模型在文化表现上的独特特性。

Comments 16 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.07835 2026-01-13 cs.CR cs.CV 80%

SecureCAI: Injection-Resilient LLM Assistants for Cybersecurity Operations

SecureCAI: 面向网络安全操作的抗注入大语言模型助手

Mohammed Himayath Ali, Mohammed Aqib Abdullah, Mohammed Mudassir Uddin, Shahnawaz Alam

机构 * Computer Science Department, Cybersecurity and Artificial Intelligence Division(计算机科学系、网络安全与人工智能 division)

专题命中 偏好对齐 :safety(abstract);prompt injection(abstract);trustworthy(abstract);AI safety(abstract)

AI总结 SecureCAI通过引入安全意识护栏和偏好优化,有效提升网络安全操作中对抗攻击的防御能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.16531 2025-09-23 cs.CL cs.AI cs.FL cs.LG 80%

Bayesian scaling laws for in-context learning

Aryaman Arora, Dan Jurafsky, Christopher Potts, Noah D. Goodman

机构 * Stanford University(斯坦福大学)

专题命中 偏好对齐 :alignment(abstract);DPO(abstract);safety(abstract);分类 cs.CL、cs.AI、cs.LG

Comments COLM 2025 camera-ready version; 9 pages main text, 39 pages total

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17828 2025-07-04 cs.LG cs.AI cs.CL 80%

Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach

Xinnan Zhang, Chenliang Li, Siliang Zeng, Jiaxiang Li, Zhongruo Wang, Kaixiang Lin, Songtao Lu, Alfredo Garcia, Mingyi Hong

专题命中 偏好对齐 :alignment(abstract);RLHF(abstract);DPO(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.09401 2025-02-12 cs.LG cs.AI cs.CL math.OC stat.ML 80%

Reinforcement Learning from Human Feedback with Active Queries

Kaixuan Ji, Jiafan He, Quanquan Gu

专题命中 偏好对齐 :alignment(abstract);RLHF(abstract);DPO(abstract);分类 cs.CL、cs.AI、cs.LG

Comments 28 pages, 1 figure, 4 table

详情

展开后加载摘要…

URL PDF HTML 收藏