arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 3266 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 偏好对齐 3266 篇

2504.16628 2026-01-01 cs.LG cs.CL 81%

ParetoHqD: Fast Offline Multiobjective Alignment of Large Language Models using Pareto High-quality Data

ParetoHqD: 利用帕累托高质量数据快速实现大语言模型的多目标对齐

Haoran Gu, Handing Wang, Yi Mei, Mengjie Zhang, Yaochu Jin

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.CL、cs.LG

AI总结 ParetoHqD通过利用帕累托高质量数据实现大语言模型的多目标对齐,提升对齐效果和效率。

Comments Accepted as a main conference paper at AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15811 2025-12-16 cs.CL cs.AI 81%

From Clicks to Preference: A Multi-stage Alignment Framework for Generative Query Suggestion in Conversational System

从点击到偏好:一种多阶段对齐框架用于对话系统中的生成查询建议

Junhao Yin, Haolin Wang, Peng Bao, Ju Xu, Yongliang Wang

机构 * Bytedance Shanghai China(字节跳动上海中国) Bytedance Beijing China(字节跳动北京中国)

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.CL、cs.AI

AI总结 本文提出了一种多阶段对齐框架,通过提示工程、知识蒸馏和高斯奖励模型,提升生成式查询建议的用户偏好对齐效果,实验显示在自动和人工评估中均优于基线,并提高用户参与度34%

Comments Accepted by SIGKDD Conference on Knowledge Discovery and Data Mining (KDD 26)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13426 2025-12-16 cs.CL cs.AI 81%

ALIGN: Word Association Learning for Cultural Alignment in Large Language Models

ALIGN: 用于大语言模型中文化对齐的词关联学习

Chunhua Liu, Kabir Manandhar Shrestha, Sukai Huang

机构 * School of Computing and Information Systems, The University of Melbourne(计算与信息系统学院,墨尔本大学) Melbourne Data Analytics Platform, The University of Melbourne(墨尔本数据分析平台,墨尔本大学) Faculty of Information Technology, Monash University(信息技术学院,墨尔本大学)

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.CL、cs.AI

AI总结 ALIGN通过微调大语言模型以学习母语使用者的词关联规范,显著提升文化对齐效果,无需昂贵重新训练。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19700 2025-12-16 cs.CL cs.AI 81%

Leveraging Importance Sampling to Detach Alignment Modules from Large Language Models

利用重要性采样将对齐模块从大语言模型中分离出来

Yi Liu, Dianqing Liu, Mingye Zhu, Junbo Guo, Yongdong Zhang, Zhendong Mao

机构 * State Key Laboratory of Communication Content Cognition, People’s Daily Online(通信内容认知国家重点实验室,人民在线) University of Science and Technology of China(中国科学技术大学)

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.CL、cs.AI

AI总结 本文提出残差对齐模型(RAM),通过重要性采样将对齐模块与大语言模型分离,提升灵活性和可扩展性,并在多种任务上优于基线模型。

Comments Accepted by NeurIPS 2025, 28 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.21101 2025-12-10 cs.CL cs.LG 81%

Mortgage Language Model: Domain-Adaptive Pretraining with Residual Instruction, Alignment Tuning, and Task-Specific Routing

抵押贷款语言模型:带有残差指令、对齐微调和任务特定路由的领域自适应预训练

Manish Jain, Satheesh Kumar Ponnambalam, Salman Faroz, Chandrakanth Lns, Vinay Sharma

机构 * Firstsource

专题命中 偏好对齐 :alignment(title);DPO(abstract);分类 cs.CL、cs.LG

AI总结 MortgageLLM通过双专家架构和残差指令技术,在抵押贷款领域实现领域自适应预训练,提升对话问答和结构化任务性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.06196 2025-12-09 cs.AI cs.CL 81%

ARCANE: A Multi-Agent Framework for Interpretable and Configurable Alignment

ARCANE:一种多智能体框架,用于可解释和可配置的对齐

Charlie Masters, Marta Grześkiewicz, Stefano V. Albrecht

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.CL、cs.AI

AI总结 ARCANE是一种多智能体框架,通过动态规则生成实现可解释和可配置的对齐,适用于复杂长期任务。

Comments Accepted to the AAAI 2026 LLAMAS Workshop (Large Language Model Agents for Multi-Agent Systems)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.19023 2025-11-25 cs.LG cs.AI 81%

OrdMoE: Preference Alignment via Hierarchical Expert Group Ranking in Multimodal Mixture-of-Experts LLMs

OrdMoE:通过多模态混合专家模型中的层次专家小组排名实现偏好对齐

Yuting Gao, Weihao Chen, Lan Wang, Ruihan Xu, Qingpei Guo

机构 * AntGroup(蚂蚁集团) Peking University(北京大学)

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.AI、cs.LG

AI总结 OrdMoE通过利用混合专家架构中的内在信号,实现多模态混合专家语言模型的零成本偏好对齐,无需人工标注数据即可提升模型对齐和性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13007 2025-11-18 cs.AI cs.LG 81%

GEM: Generative Entropy-Guided Preference Modeling for Few-shot Alignment of LLMs

Yiyang Zhao, Huiyu Bai, Xuejiao Zhao

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.AI、cs.LG

Comments This paper has been accepted by AAAI 2026-AIA and designated as an oral presentation paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12573 2025-11-18 cs.CL cs.AI 81%

Mitigating Length Bias in RLHF through a Causal Lens

Hyeonji Kim, Sujeong Oh, Sanghack Lee

专题命中 偏好对齐 :RLHF(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.06965 2025-11-18 cs.CL cs.AI 81%

Uncovering Factor Level Preferences to Improve Human-Model Alignment

Juhyun Oh, Eunsu Kim, Jiseon Kim, Wenda Xu, Inha Cha, William Yang Wang, Alice Oh

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10656 2025-11-17 cs.CL cs.AI 81%

Preference Orchestrator: Prompt-Aware Multi-Objective Alignment for Large Language Models

Biao Liu, Ning Xu, Junming Yang, Xin Geng

机构 * School of Computer Science and Engineering(计算机科学与工程学院) Key Laboratory of New Generation Artificial Intelligence Technology and Its Interdisciplinary Applications(新一代人工智能技术及其交叉应用关键实验室) Ministry of Education(教育部)

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05359 2025-11-10 cs.CR cs.CL cs.CY 81%

ConVerse: Benchmarking Contextual Safety in Agent-to-Agent Conversations

Amr Gomaa, Ahmed Salem, Sahar Abdelnabi

机构 * German Research Center for Artificial Intelligence (DFKI)(德国人工智能研究中心(DFKI)) Microsoft(微软) ELLIS Institute Tübingen and MPI for Intelligent Systems(图宾根ELLIS研究所和智能系统研究所) Tübingen AI Center(图宾根人工智能中心)

专题命中 偏好对齐 :safety(title,abstract);分类 cs.CL、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01183 2025-10-30 cs.LG cs.AI stat.ML 81%

Doubly Robust Alignment for Large Language Models

Erhan Xu, Kai Ye, Hongyi Zhou, Luhan Zhu, Francesco Quinzan, Chengchun Shi

机构 * Department of Statistics(统计系) LSE London, UK(伦敦大学学院) Department of Mathematics(数学系) Tsinghua University(清华大学) School of Design LCC, UAL London, UK(伦敦艺术大学设计学院) Department of Engineering Science(工程科学系) University of Oxford(牛津大学)

专题命中 偏好对齐 :alignment(title);RLHF(abstract);分类 cs.AI、cs.LG

Comments Accepted to NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24700 2025-10-29 cs.LG cs.AI cs.IT math.IT stat.ML 81%

Greedy Sampling Is Provably Efficient for RLHF

Di Wu, Chengshuai Shi, Jing Yang, Cong Shen

机构 * Electrical and Computer Engineering University of Virginia(电气与计算机工程大学弗吉尼亚大学) Princeton Language and Intelligence Princeton University(普林斯顿语言与智能普林斯顿大学)

专题命中 偏好对齐 :RLHF(title,abstract);分类 cs.AI、cs.LG

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21798 2025-10-27 cs.CL cs.AI 81%

Evaluating and Improving Cultural Awareness of Reward Models for LLM Alignment

Hongbin Zhang, Kehai Chen, Xuefeng Bai, Yang Xiang, Min Zhang

机构 * Institute of Computing and Intelligence, Harbin Institute of Technology, Shenzhen, China(计算与智能研究所,哈尔滨工业大学,深圳,中国) Peng Cheng Laboratory, Shenzhen, China(鹏城实验室,深圳,中国)

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.CL、cs.AI

Comments Under review;Work in progress;

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.17041 2025-10-21 cs.CV cs.AI cs.LG 81%

Free$^2$Guide: Training-Free Text-to-Video Alignment using Image LVLM

Jaemin Kim, Bryan Sangwoo Kim, Jong Chul Ye

机构 * Graduate School of AI, KAIST(人工智能研究生院,韩国科学技术院)

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.AI、cs.LG

Comments ICCV 2025 accepted

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23564 2025-10-15 cs.AI cs.CL 81%

Clean First, Align Later: Benchmarking Preference Data Cleaning for Reliable LLM Alignment

Samuel Yeh, Sharon Li

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.CL、cs.AI

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01700 2025-10-03 cs.AI cs.CV cs.LG 81%

VaPR -- Vision-language Preference alignment for Reasoning

Rohan Wadhawan, Fabrice Y Harel-Canada, Zi-Yi Dou, Suhaila Shakiah, Robinson Piramuthu, Nanyun Peng

机构 * Department of Computer Science, University of California Los Angeles(加州大学洛杉矶分校计算机科学系) Amazon.com, Inc.(亚马逊公司)

专题命中 偏好对齐 :alignment(title);DPO(abstract);分类 cs.AI、cs.LG

Journal ref COLM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24713 2025-09-30 cs.LG cs.AI 81%

Circuit-Aware Reward Training: A Mechanistic Framework for Longtail Robustness in RLHF

Jing Liu

机构 * ENS, Université PSL, EHESS, CNRS(高等科学研究所、巴黎国立大学、EHESS、国家科学研究中心)

专题命中 偏好对齐 :RLHF(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.05381 2025-09-16 cs.AI cs.LG 81%

Murphys Laws of AI Alignment: Why the Gap Always Wins

Madhava Gaikwad

机构 * Microsoft(微软)

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.AI、cs.LG

Comments Provides a formal impossibility theorem (Murphys Gap) and welcomes collaboration on large-scale experiments and benchmark design

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.10717 2025-08-27 cs.CL cs.AI 81%

A Modular Approach for Clinical SLMs Driven by Synthetic Data with Pre-Instruction Tuning, Model Merging, and Clinical-Tasks Alignment

Jean-Philippe Corbeil, Amin Dada, Jean-Michel Attendu, Asma Ben Abacha, Alessandro Sordoni, Lucas Caccia, François Beaulieu, Thomas Lin, Jens Kleesiek, Paul Vozila

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.CL、cs.AI

Journal ref ACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17000 2025-08-26 cs.CL cs.LG 81%

KL-Regularised Q-Learning: A Token-level Action-Value perspective on Online RLHF

Jason R Brown, Lennie Wells, Edward James Young, Sergio Bacallado

机构 * Computational and Biological Learning Group, Department of Engineering, University of Cambridge, Cambridge, UK(计算生物学学习组,工程系,剑桥大学,剑桥,英国) Department of Computer Science and Technology, University of Cambridge, Cambridge, UK(计算机科学与技术系,剑桥大学,剑桥,英国) Statistics Laboratory, Department of Pure Mathematics and Mathematical Statistics, University of Cambridge, UK(统计实验室,纯粹数学与数学统计系,剑桥大学,英国)

专题命中 偏好对齐 :RLHF(title,abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01674 2025-08-08 cs.CL cs.AI cs.HC 81%

CUPID: Evaluating Personalized and Contextualized Alignment of LLMs from Interactions

Tae Soo Kim, Yoonjoo Lee, Yoonah Park, Jiho Kim, Young-Ho Kim, Juho Kim

机构 * KAIST(韩国科学技术院) Seoul National University(首尔国立大学) Calvin University(凯尔文大学) NAVER AI LAB(NAVER AI实验室)

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.CL、cs.AI

Comments Accepted to COLM 2025. Project Website: https://cupid.kixlab.org/

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04626 2025-08-07 cs.CL cs.AI 81%

P-Aligner: Enabling Pre-Alignment of Language Models via Principled Instruction Synthesis

Feifan Song, Bofei Gao, Yifan Song, Yi Liu, Weimin Xiong, Yuyang Song, Tianyu Liu, Guoyin Wang, Houfeng Wang

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01930 2025-08-05 cs.CL cs.AI 81%

Word Overuse and Alignment in Large Language Models: The Influence of Learning from Human Feedback

Tom S. Juzek, Zina B. Ward

机构 * Florida State University(佛罗里达州立大学)

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.CL、cs.AI

Comments Accepted for publication in the Proceedings of the 5th Workshop on Bias and Fairness in AI (BIAS 2025) at ECML PKDD

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22789 2025-08-01 cs.LG cs.AI 81%

G-Core: A Simple, Scalable and Balanced RLHF Trainer

Junyu Wu, Weiming Chang, Xiaotao Liu, Guanyou He, Haoqiang Hong, Boqi Liu, Hongtao Tian, Tao Yang, Yunsheng Shi, Feng Lin, Ting Yao

专题命中 偏好对齐 :RLHF(title,abstract);分类 cs.AI、cs.LG

Comments I haven't received company approval yet, and I uploaded it by mistake

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20335 2025-07-29 cs.LG cs.AI 81%

Cultivating Helpful, Personalized, and Creative AI Tutors: A Framework for Pedagogical Alignment using Reinforcement Learning

Siyu Song, Wentao Liu, Ye Lu, Ruohua Zhang, Tao Liu, Jinze Lv, Xinyun Wang, Aimin Zhou, Fei Tan, Bo Jiang, Hao Hao

机构 * Shanghai Innavation Institute(上海创新研究院) Shanghai Institute of AI for Education(上海人工智能教育研究院) School of Computer Science and Technology(计算机科学与技术学院) Department of Educational Information Technology(教育信息技术系)

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.05236 2025-07-24 cs.SD cs.AI cs.LG eess.AS 81%

Koel-TTS: Enhancing LLM based Speech Generation with Preference Alignment and Classifier Free Guidance

Shehzeen Hussain, Paarth Neekhara, Xuesong Yang, Edresson Casanova, Subhankar Ghosh, Mikyas T. Desta, Roy Fejgin, Rafael Valle, Jason Li

机构 * NVIDIA Corporation(英伟达公司)

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.AI、cs.LG

Journal ref ICML Workshop on Machine Learning for Audio, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.17141 2025-07-22 cs.CL cs.AI 81%

MetaAligner: Towards Generalizable Multi-Objective Alignment of Language Models

Kailai Yang, Zhiwei Liu, Qianqian Xie, Jimin Huang, Tianlin Zhang, Sophia Ananiadou

机构 * The University of Manchester(曼彻斯特大学)

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.CL、cs.AI

Comments Accepted by NeurIPS 2024 main track

Journal ref https://proceedings.neurips.cc/paper_files/paper/2024/hash/3d03800841fa1bb2f43ef1750aafcce4-Abstract-Conference.html

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.02882 2025-07-15 cs.CL cs.LG 81%

DiaTool-DPO: Multi-Turn Direct Preference Optimization for Tool-Augmented Large Language Models

Sunghee Jung, Donghun Lee, Shinbok Lee, Gaeun Seo, Daniel Lee, Byeongil Ko, Junrae Cho, Kihyun Kim, Eunggyun Kim, Myeongcheol Shin

机构 * Kakao Corp.(韩国 Kakao 公司)

专题命中 偏好对齐 :DPO(title,abstract);分类 cs.CL、cs.LG

Comments Accepted to SIGDIAL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏