arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 3266 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 偏好对齐 3266 篇

2406.09574 2025-05-19 cs.LG 79%

Online Bandit Learning with Offline Preference Data for Improved RLHF

Akhil Agnihotri, Rahul Jain, Deepak Ramachandran, Zheng Wen

机构 * University of Southern California(南加州大学) Google DeepMind(谷歌DeepMind) USC(南加州大学)

专题命中 偏好对齐 :RLHF(title,abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.03059 2025-05-07 cs.CL 79%

Improving Model Alignment Through Collective Intelligence of Open-Source LLMS

Junlin Wang, Roy Xie, Shang Zhu, Jue Wang, Ben Athiwaratkun, Bhuwan Dhingra, Shuaiwen Leon Song, Ce Zhang, James Zou

机构 * Duke University(杜克大学) Together AI University of Chicago(芝加哥大学) Stanford University(斯坦福大学)

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.CL

Comments ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.17003 2025-05-06 cs.CL 79%

A Survey on Personalized Alignment -- The Missing Piece for Large Language Models in Real-World Applications

Jian Guan, Junfei Wu, Jia-Nan Li, Chuanqi Cheng, Wei Wu

机构 * Ant Group(蚂蚁集团) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学人工智能学院)

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.CL

Comments Survey paper; 11 pages; Literature reviewed up to ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.01957 2025-05-02 cs.CL 79%

Challenges and Future Directions of Data-Centric AI Alignment

Min-Hsuan Yeh, Jeffrey Wang, Xuefeng Du, Seongheon Park, Leitian Tao, Shawn Im, Yixuan Li

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.CL

Comments ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.08813 2025-04-28 cs.CL 79%

Your Weak LLM is Secretly a Strong Teacher for Alignment

Leitian Tao, Yixuan Li

机构 * Department of Computer Sciences, University of Wisconsin-Madison(计算机科学系,威斯康星大学麦迪逊分校)

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.CL

Comments Accepted by ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.15719 2025-04-23 cs.AI 79%

Implementing Rational Choice Functions with LLMs and Measuring their Alignment with User Preferences

Anna Karnysheva, Christian Drescher, Dietrich Klakow

机构 * Spoken Language Systems (LSV)(语音语言系统(LSV)) Saarland University(萨尔兰大学) Mercedes-Benz AG(梅赛德斯-奔驰集团)

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.14148 2025-04-22 cs.CV cs.CL 79%

Fine-Grained Verifiers: Preference Modeling as Next-token Prediction in Vision-Language Alignment

Chenhang Cui, An Zhang, Yiyang Zhou, Zhaorun Chen, Gelei Deng, Huaxiu Yao, Tat-Seng Chua

机构 * National University of Singapore(新加坡国立大学) UNC-Chapel Hill(北卡罗来纳大学教堂山分校) University of Chicago(芝加哥大学) Nanyang Technological University(南洋理工大学)

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.CL

Comments 23 pages; Published as a conference paper at ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.07070 2025-04-10 cs.CL 79%

A Survey on Personalized and Pluralistic Preference Alignment in Large Language Models

Zhouhang Xie, Junda Wu, Yiran Shen, Yu Xia, Xintong Li, Aaron Chang, Ryan Rossi, Sachin Kumar, Bodhisattwa Prasad Majumder, Jingbo Shang, Prithviraj Ammanabrolu, Julian McAuley

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.04657 2025-04-08 cs.LG 79%

ACE-RLHF: Automated Code Evaluation and Socratic Feedback Generation Tool using Large Language Models and Reinforcement Learning with Human Feedback

Tasnia Rahman, Sathish A. P. Kumar, Sumit Jha, Arvind Ramanathan

专题命中 偏好对齐 :RLHF(title,abstract);分类 cs.LG

Comments 9 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.09893 2025-04-07 cs.CL 79%

RMB: Comprehensively Benchmarking Reward Models in LLM Alignment

Enyu Zhou, Guodong Zheng, Binghai Wang, Zhiheng Xi, Shihan Dou, Rong Bao, Wei Shen, Limao Xiong, Jessica Fan, Yurong Mou, Rui Zheng, Tao Gui, Qi Zhang, Xuanjing Huang

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.CL

Comments Accepted by ICLR2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.20105 2025-03-27 cs.AI cs.RO 79%

Direct Post-Training Preference Alignment for Multi-Agent Motion Generation Models Using Implicit Feedback from Pre-training Demonstrations

Ran Tian, Kratarth Goel

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.AI

Comments ICLR 2025 Spotlight

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.16094 2025-03-21 cs.CL 79%

Cultural Alignment in Large Language Models Using Soft Prompt Tuning

Reem I. Masoud, Martin Ferianc, Philip Treleaven, Miguel Rodrigues

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.08035 2025-03-12 cs.CL 79%

Group Preference Alignment: Customized LLM Response Generation from In-Situ Conversations

Ishani Mondal, Jack W. Stokes, Sujay Kumar Jauhar, Longqi Yang, Mengting Wan, Xiaofeng Xu, Xia Song, Jennifer Neville

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.CL

Comments 23 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.04858 2025-03-10 cs.CV cs.AI 79%

SHAPE : Self-Improved Visual Preference Alignment by Iteratively Generating Holistic Winner

Kejia Chen, Jiawen Zhang, Jiacong Hu, Jiazhen Yang, Jian Lou, Zunlei Feng, Mingli Song

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.17927 2025-03-06 cs.CL 79%

Advantage-Guided Distillation for Preference Alignment in Small Language Models

Shiping Gao, Fanqi Wan, Jiajian Guo, Xiaojun Quan, Qifan Wang

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.CL

Comments Accepted by ICLR 2025(spotlight)

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.11817 2025-03-04 cs.CV cs.LG cs.MM 79%

Improving Long-Text Alignment for Text-to-Image Diffusion Models

Luping Liu, Chao Du, Tianyu Pang, Zehan Wang, Chongxuan Li, Dong Xu

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.LG

Journal ref International Conference on Learning Representations (ICLR 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.20757 2025-03-03 cs.CL 79%

The Rise of Darkness: Safety-Utility Trade-Offs in Role-Playing Dialogue Agents

Yihong Tang, Kehai Chen, Xuefeng Bai, Zhengyu Niu, Bo Wang, Jie Liu, Min Zhang

专题命中 偏好对齐 :safety(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.14356 2025-02-21 cs.CL 79%

Full-Step-DPO: Self-Supervised Preference Optimization with Step-wise Rewards for Mathematical Reasoning

Huimin Xu, Xin Mao, Feng-Lin Li, Xiaobao Wu, Wang Chen, Wei Zhang, Anh Tuan Luu

专题命中 偏好对齐 :DPO(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.06812 2025-02-19 cs.LG cs.GR 79%

Harness Local Rewards for Global Benefits: Effective Text-to-Video Generation Alignment with Patch-level Reward Models

Shuting Wang, Haihong Tang, Zhicheng Dou, Chenyan Xiong

专题命中 偏好对齐 :alignment(title);DPO(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.02743 2025-02-18 cs.CL 79%

MA-RLHF: Reinforcement Learning from Human Feedback with Macro Actions

Yekun Chai, Haoran Sun, Huang Fang, Shuohuan Wang, Yu Sun, Hua Wu

专题命中 偏好对齐 :RLHF(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.11107 2025-02-12 cs.CL 79%

Exploring Safety-Utility Trade-Offs in Personalized Language Models

Anvesh Rao Vijjini, Somnath Basu Roy Chowdhury, Snigdha Chaturvedi

专题命中 偏好对齐 :safety(title,abstract);分类 cs.CL

Comments NAACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.10771 2025-02-04 cs.LG stat.ML 79%

A Probabilistic Approach for Model Alignment with Human Comparisons

Junyu Cao, Mohsen Bayati

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.04185 2025-01-09 cs.CL 79%

HAF-RM: A Hybrid Alignment Framework for Reward Model Training

Shujun Liu, Xiaoyu Shen, Yuhang Lai, Siyuan Wang, Shengbin Yue, Zengfeng Huang, Xuanjing Huang, Zhongyu Wei

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.10616 2024-12-17 cs.LG 79%

Hybrid Preference Optimization for Alignment: Provably Faster Convergence Rates by Combining Offline Preferences with Online Exploration

Avinandan Bose, Zhihan Xiong, Aadirupa Saha, Simon Shaolei Du, Maryam Fazel

专题命中 偏好对齐 :alignment(title);RLHF(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.12822 2024-12-10 cs.CL 79%

Language Models Learn to Mislead Humans via RLHF

Jiaxin Wen, Ruiqi Zhong, Akbir Khan, Ethan Perez, Jacob Steinhardt, Minlie Huang, Samuel R. Bowman, He He, Shi Feng

专题命中 偏好对齐 :RLHF(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.16019 2024-12-04 cs.CL 79%

The PRISM Alignment Dataset: What Participatory, Representative and Individualised Human Feedback Reveals About the Subjective and Multicultural Alignment of Large Language Models

Hannah Rose Kirk, Alexander Whitefield, Paul Röttger, Andrew Bean, Katerina Margatina, Juan Ciro, Rafael Mosquera, Max Bartolo, Adina Williams, He He, Bertie Vidgen, Scott A. Hale

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.CL

Journal ref The Thirty-eight Conference on Neural Information Processing Systems Datasets and Benchmarks Track (2024)

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.15193 2024-11-27 cs.CL 79%

Inference Time Alignment with Reward-Guided Tree Search

Chia-Yu Hung, Navonil Majumder, Ambuj Mehrish, Soujanya Poria

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.07571 2024-11-18 cs.CL cs.CV 79%

How Does Vision-Language Adaptation Impact the Safety of Vision Language Models?

Seongyun Lee, Geewook Kim, Jiyeon Kim, Hyunji Lee, Hoyeon Chang, Sue Hyun Park, Minjoon Seo

专题命中 偏好对齐 :safety(title,abstract);分类 cs.CL

Comments Work in Progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.09341 2024-11-15 cs.LG 79%

Approximated Variational Bayesian Inverse Reinforcement Learning for Large Language Model Alignment

Yuang Cai, Yuyu Yuan, Jinsheng Shi, Qinhong Lin

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.02712 2024-11-06 cs.CV cs.AI 79%

V-DPO: Mitigating Hallucination in Large Vision Language Models via Vision-Guided Direct Preference Optimization

Yuxi Xie, Guanzhen Li, Xiao Xu, Min-Yen Kan

专题命中 偏好对齐 :DPO(title,abstract);分类 cs.AI

Comments EMNLP 2024 Findings; 9 pages, 6 figures, 5 tables (16 pages, 8 figures, 8 tables including references and appendices)

详情

展开后加载摘要…

URL PDF HTML 收藏