arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 3266 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 偏好对齐 3266 篇

2412.01253 2025-01-23 cs.CL cs.AI cs.LG 80%

Yi-Lightning Technical Report

Alan Wake, Bei Chen, C. X. Lv, Chao Li, Chengen Huang, Chenglin Cai, Chujie Zheng, Daniel Cooper, Fan Zhou, Feng Hu, Ge Zhang, Guoyin Wang, Heng Ji, Howard Qiu, Jiangcheng Zhu, Jun Tian, Katherine Su, Lihuan Zhang, Liying Li, Ming Song, Mou Li, Peng Liu, Qicheng Hu, Shawn Wang, Shijun Zhou, Shiming Yang, Shiyong Li, Tianhang Zhu, Wen Xie, Wenhao Huang, Xiang He, Xiaobo Chen, Xiaohui Hu, Xiaoyi Ren, Xinyao Niu, Yanpeng Li, Yongke Zhao, Yongzhen Luo, Yuchi Xu, Yuxuan Sha, Zhaodong Yan, Zhiyuan Liu, Zirui Zhang, Zonghong Dai

专题命中 偏好对齐 :RLHF(abstract);safety(abstract);AI safety(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.05857 2025-01-07 cs.CL cs.AI cs.LG 80%

Improving Summarization with Human Edits

Zonghai Yao, Benjamin J Schloss, Sai P. Selvaraj

专题命中 偏好对齐 :alignment(abstract);RLHF(abstract);DPO(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Proceedings of the Main Conference on Empirical Methods in Natural Language Processing (EMNLP) 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.19534 2024-11-01 cs.LG cs.AI cs.CL 80%

Preference Learning Algorithms Do Not Learn Preference Rankings

Angelica Chen, Sadhika Malladi, Lily H. Zhang, Xinyi Chen, Qiuyi Zhang, Rajesh Ranganath, Kyunghyun Cho

专题命中 偏好对齐 :alignment(abstract);RLHF(abstract);DPO(abstract);分类 cs.CL、cs.AI、cs.LG

Comments NeurIPS 2024 camera-ready

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.13542 2024-07-19 cs.CL cs.AI cs.LG 80%

Self-play with Execution Feedback: Improving Instruction-following Capabilities of Large Language Models

Guanting Dong, Keming Lu, Chengpeng Li, Tingyu Xia, Bowen Yu, Chang Zhou, Jingren Zhou

专题命中 偏好对齐 :alignment(abstract);RLHF(abstract);DPO(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Work in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.02479 2024-06-11 cs.LG cs.AI cs.CL cs.HC 80%

BRAIn: Bayesian Reward-conditioned Amortized Inference for natural language generation from feedback

Gaurav Pandey, Yatin Nandwani, Tahira Naseem, Mayank Mishra, Guangxuan Xu, Dinesh Raghu, Sachindra Joshi, Asim Munawar, Ramón Fernandez Astudillo

专题命中 偏好对齐 :alignment(abstract);RLHF(abstract);DPO(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Accepted at ICML 2024 (main conference)

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.17546 2024-05-01 cs.LG cs.AI cs.CL stat.ML 80%

Probabilistic Inference in Language Models via Twisted Sequential Monte Carlo

Stephen Zhao, Rob Brekelmans, Alireza Makhzani, Roger Grosse

专题命中 偏好对齐 :RLHF(abstract);safety(abstract);harmlessness(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.10683 2024-02-20 cs.CL cs.AI cs.LG 80%

Large Language Model Unlearning

Yuanshun Yao, Xiaojun Xu, Yang Liu

专题命中 偏好对齐 :alignment(abstract);RLHF(abstract);red teaming(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.16763 2023-10-26 cs.CL cs.AI cs.LG 80%

SuperHF: Supervised Iterative Learning from Human Feedback

Gabriel Mukobi, Peter Chatain, Su Fong, Robert Windesheim, Gitta Kutyniok, Kush Bhatia, Silas Alberti

专题命中 偏好对齐 :alignment(abstract);RLHF(abstract);safety(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Accepted to the Socially Responsible Language Modelling Research (SoLaR) workshop at NeurIPS 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.12845 2024-06-19 cs.LG cs.CL 80%

Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts

Haoxiang Wang, Wei Xiong, Tengyang Xie, Han Zhao, Tong Zhang

专题命中 偏好对齐 :RLHF(abstract,comments);alignment(abstract);safety(abstract);分类 cs.CL、cs.LG

Comments Technical report v1. Code and model are released at https://github.com/RLHFlow/RLHF-Reward-Modeling/

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.24046 2026-08-26 cs.AI 新提交 79%

Algorithmic Impact Reveals the Hidden Social Choice Structure of Alignment

算法影响揭示对齐的隐藏社会选择结构

Zachary Wojtowicz, Michelle Si, Finale Doshi-Velez, Ariel Procaccia

机构 * MIT(麻省理工学院) Harvard University(哈佛大学)

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.AI

AI总结 该研究将AI对齐问题转化为凸影响空间上的线性优化,结合福利经济学与机制设计,推导了防策略的社会选择机制及最大化功利主义社会福利的对齐协议,并通过多领域人类偏好实证验证其福利影响。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.01160 2026-08-25 cs.AI 版本更新 79%

Expected Value Alignment for Generative Reward Modeling in Formal Mathematics Verification

形式数学验证中生成式奖励建模的期望值对齐

Shihao Ji, Haotao Tan, Zihui Song, Mingyu Li

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.AI

AI总结 提出期望值对齐(EVA)方法,通过从模型词元分布中提取连续分数,在保持生成式奖励模型离散输出的同时实现连续评分,用于Lean 4形式验证。

Comments Withdrawn due to serious concerns regarding the authenticity and accuracy of the listed authorship. The identity of one or more listed authors cannot presently be verified, and the author list may not represent distinct contributors. The manuscript is withdrawn pending institutional review

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.20512 2026-08-24 cs.CY 新提交 79%

From Urban Mobility to Epidemic Dynamics: A Mixture-of-Experts Framework with Preference Alignment for Policy Scenario Simulation

从城市流动性到流行病动力学:用于政策情景模拟的带偏好对齐的混合专家框架

Yun Ye, Arsalan Dezhkam, Junyuan Liu, Xinglei Wang, Tao Cheng

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.CY

AI总结 该研究提出带偏好对齐的混合专家框架UrbanShare-MoE-PA,通过智能体级流动性建模连接政策、行为与流行病模拟,评估了不同封锁政策的流行病-活动权衡,为NPI情景分析提供了可解释工具。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.18132 2026-08-20 cs.CL cs.SD eess.AS 新提交 79%

Alignment Is All You Need: Instruction-Free Training for General Audio-Language Models

对齐即全部所需:通用音频-语言模型的无指令训练

Xuanru Zhou, Yiwen Shao, Jiahong Li, Dong Yu

机构 * Zhejiang University(浙江大学) Tencent Hunyuan(腾讯混元)

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.CL

AI总结 该研究提出仅对齐的无指令大音频-语言模型LALM,仅训练轻量投影器,在多模态数据集上用更少数据达到或优于基线,证明仅靠对齐即可构建有竞争力的多模态大语言模型。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.27366 2026-08-18 cs.CL 版本更新 79%

BridgeAlign: Bridging Preference Alignment for Humanities and Social Sciences

BridgeAlign:面向人文社科的偏好对齐框架

Ru Peng, Haokai Xu, Xijun Gu, Tianyu Zhao, Zhiting Fan, Yawen Zeng, Yihong Zhuang, Jinyang Zhang, Kexin Yang, Jian Wu, Hao Chen, Junyang Lin, Dayiheng Liu, Junbo Zhao

机构 * Alibaba Group(阿里巴巴集团) Ant Group(蚂蚁集团)

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.CL

AI总结 该研究针对人文社科领域的偏好对齐需求,提出BridgeAlign框架,经21万+合成偏好样本对齐后,使Qwen3-8B在17个基准测试中优于11个强基线,且人工偏好与知识能力无权衡。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.09164 2026-08-14 cs.AI 版本更新 79%

CIDER: A Dataset of Contextual Disclosure Boundaries for Privacy Preference Alignment

CIDER:用于隐私偏好对齐的上下文披露边界数据集

Bingcan Guo, Eryue Xu, Jijie Zhou, Zhiping Zhang, Tianshi Li

机构 * University of Washington(华盛顿大学) UIUC(伊利诺伊大学厄巴纳-香槟分校) Northeastern University(东北大学)

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.AI

AI总结 本文推出包含169名用户标注的CIDER数据集,用于评估LLM隐私偏好对齐,发现上下文个性化可提升预测准确率,GPT-5.4和Claude Sonnet 4.6表现更优。

Comments Accepted to COLM 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.06792 2026-08-10 cs.IR cs.AI 新提交 79%

Progressive Alignment of Recommender Foundation Model through Multi-Phase Post-Training

通过多阶段后训练实现推荐系统基础模型的渐进式对齐

Oseong Choi, Hoeinn Kim, Jihoon Lee, Byungsoo Kang, Taeyeong Jang

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.AI

AI总结 本文提出LP-FFT-RFT三阶段渐进式后训练框架,将推荐系统基础模型的下游适配与业务指标对齐分离,实验证实其能提升推荐质量。

Comments 9 pages, 3 figures. Accepted to the 20th ACM Conference on Recommender Systems (RecSys '26), Industry Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.13227 2026-08-07 cs.CL 版本更新 79%

PolyAlign: Conditional Human-Distribution Alignment

PolyAlign: 条件性人类分布对齐

L. D. M. S. Sai Teja, Ufaq Khan, Sathira Silva, Xiao Wu, Muhammad Haris Khan

机构 * NIT Silchar(印度国立理工学院锡尔恰尔分校) MBZUAI(穆罕默德·本·扎耶德人工智能大学)

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.CL

AI总结 提出PolyAlign框架,通过桶感知SFT和人类分布偏好优化,实现语言模型在不同交互上下文中的条件性人类分布对齐,提升自然性和分布忠实度。

Comments 23 pages, 5 Figures, 13 Tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.02171 2026-08-04 cs.AI 新提交 79%

From Profiling to Synthesis: Benchmarking Implicit Behavioral Alignment in Personalized LLM Agents

从分析到综合:个性化大语言模型智能体中的隐式行为对齐基准测试

Jiajia Song, Bobo Li, Haiwen Yi, Zibo Ji, Meishan Zhang, Hao Fei, Min Zhang, Mong-Li Lee, Wynne Hsu

机构 * National University of Singapore(新加坡国立大学) University of Toronto(多伦多大学) University of Minnesota Twin Cities(明尼苏达大学双城分校) Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳)) University of Oxford(牛津大学)

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.AI

AI总结 该研究针对大语言模型智能体的个性化问题,构建了IBA-Bench基准,提出IBA-Agent框架,实验显示其可在九类场景中提升隐式行为对齐效果,但有效个性化仍是重大挑战。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.14990 2026-07-29 cs.AI 版本更新 79%

The Possibility of Artificial Intelligence Becoming a Subject and the Alignment Problem

人工智能成为主体的可能性及对齐问题

Till Mossakowski, Helena Esther Grass

机构 * Institute of computer science, Osnabrück University(奥斯纳布吕克大学计算机科学学院) Institute of philosophy, Oldenburg University(奥尔登堡大学哲学学院)

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.AI

AI总结 本文探讨AGI可能成为独立主体的可能性,提出通过支持自主性的养育方式替代传统控制策略,强调人类与AGI的协作共存与共进化。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.13315 2026-07-23 cs.CL 版本更新 79%

Meta-Learning Preferences for Multilingual LLM Alignment

多语言语言模型对齐的元学习偏好

Jiaying Lin, Seongho Son, Nam Phuong Tran, Long Tran-thanh, Ilija Bogunovic, Debmalya Mandal

机构 * University of Warwick(华威大学) University College London(伦敦大学学院) University of Basel(巴塞尔大学)

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.CL

AI总结 针对多语言环境下低资源语言对齐数据不足问题,提出元学习框架,利用其他语言偏好数据学习可转移初始化,经理论和实验验证,该方法在极低资源设置下比基线方法胜率最多提高28%,且在多场景表现优异。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.21550 2026-06-23 cs.AI 新提交 79%

AI Alignment From Social Choice Perspectives

从社会选择视角看AI对齐

Daniel Halpern, Evi Micha, Ariel D. Procaccia, Benjamin Schiffer, Itai Shapira, Shirley Zhang

机构 * Google Research(谷歌研究院) University of Southern California(南加州大学) Harvard University(哈佛大学)

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.AI

AI总结 本文从社会选择理论视角审视人类反馈对齐中的偏好聚合问题,识别失败模式并揭示处理分歧的广阔设计空间。

Comments Accepted for publication in ACM SIGecom Exchanges

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.13312 2026-06-23 cs.MM cs.LG 版本更新 79%

Design-MLLM: A Reinforcement Alignment Framework for Verifiable and Aesthetic Interior Design

Design-MLLM:一种用于可验证且美观的室内设计的强化对齐框架

Yuxuan Yang, Xiaotong Mao, Jingyao Wang

机构 * National Jiangsu University of Finance(江苏财经大学) University of Lorraine(洛林大学) Institute of Electronics and Information Technology, Chinese Academy of Sciences(中国科学院电子信息技术研究所) Tsinghua University(清华大学)

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.LG

AI总结 提出Design-MLLM框架,通过双分支美学导向奖励的强化对齐,解决室内设计中空间可行性硬约束与美学偏好软约束的矛盾,生成既可行又美观的设计。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25148 2026-06-19 cs.AI 版本更新 79%

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models

AAPA:用于大型语言模型后训练的对抗锚定偏好对齐

Faqiang Qian, Kang An, Weikun Zhang, Ziliang Wang, Xuhui Zheng, Liangjian Wen, Yong Dai, Mengya Gao, Yichao Wu

机构 * Southwest University of Finance and Economics(西南财经大学)

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.AI

AI总结 提出AAPA框架,通过固定轻量判别器对策略输出与专家响应进行句子级对抗锚定,增强SFT、GRPO等后训练目标,在指令遵循基准上持续提升性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.15572 2026-06-16 cs.CY 新提交 79%

Improving Capstone Team Outcomes through Dynamic Skill Matching and Preference Alignment

通过动态技能匹配和偏好对齐改进顶点项目团队成果

Brandon Pardi, Garret Castro, Michael Pisman, Avash Adhikari, Santosh Chandrasekhar

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.CY

AI总结 提出三阶段方法,结合学生偏好与项目技能需求,利用大语言模型提取技能要求,通过动态分配算法优化团队组成,提升技能覆盖和偏好满意度。

Comments 15 pages, 3 figures, Accepted to the 12th International Conference on Computational Science and Computational Intelligence (CSCI 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.14691 2026-06-15 cs.CL 新提交 79%

CORA: Analyzing and bridging thinking-answer gap in Multimodal RLVR via Consistency-Oriented Reasoning Alignment

CORA: 通过一致性导向的推理对齐分析与弥合多模态RLVR中的思考-答案差距

Jiayue Cao, Zhicong Lu, Xuehan Sun, Wei Jia, Hongling Zheng, Changyuan Tian, Zichuan Lin, Wenqian Lv, Nayu Liu

机构 * University of Chinese Academy of Sciences(中国科学院大学) Wuhan University(武汉大学) Tsinghua University(清华大学) Tianjin University(天津大学)

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.CL

AI总结 本文分析多模态RLVR中思考与答案的语义不一致问题,提出CORA方法,通过轻量级一致性奖励模型引入语义一致性,并采用混合奖励优势分裂稳定优化,提升推理忠实度。

Comments Submitted to EMNLP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.08131 2026-06-09 cs.HC cs.AI 新提交 79%

LCAM: A Framework for Diagnosing Interactional Alignment Failures in Con-versational AI

LCAM:诊断对话式AI中交互对齐失败的框架

Manuele Reani, Hongyu Tian

机构 * School of Management and Economics, The Chinese University of Hong Kong, Shenzhen(香港中文大学深圳校区管理学院)

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.AI

AI总结 提出分层认知对齐模型(LCAM),通过五层对齐和两种失调极性诊断对话式AI的交互失败,应用于LLM咨询案例揭示潜在危害。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.07877 2026-06-09 cs.CL 新提交 79%

Whose Norms? Disentangling Cultural and Personal Alignment in Large Language Models

谁的规范?解开大语言模型中的文化与个人对齐

Angana Borah, Isabelle Augenstein, Rada Mihalcea

机构 * University of Michigan - Ann Arbor(密歇根大学安娜堡分校) University of Copenhagen(哥本哈根大学)

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.CL

AI总结 提出PACT框架评估大语言模型在文化规范与个人偏好间的权衡,发现模型受国家背景影响大于年龄和性别,且人类对齐未能捕捉文化多元性。

Comments Preprint under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.01642 2026-06-08 cs.LG 版本更新 79%

Adaptive Pluralistic Alignment: A pipeline for dynamic artificial democracy

自适应多元对齐:动态人工民主的流水线

Rachel Freedman

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.LG

AI总结 提出自适应多元对齐(APA)流水线,通过低秩奖励基分解和陪审团投票机制,动态追踪社会价值观演变,避免价值锁定,无需重复预训练或大规模数据收集。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.03043 2026-06-03 cs.CL 79%

The Geometry of LLM-as-Judge: Why Inter-LLM Consensus Is Not Human Alignment

LLM作为评判者的几何学:为什么LLM间共识不等于人类对齐

Sourabrata Mukherjee, Hamna Hamna, Kalika Bali, Sunayana Sitaram

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.CL

AI总结 通过几何量测量,发现LLM评判者之间高度一致但与人类对齐差,其评分子空间与人类子空间几乎正交,共识源于子空间坍缩而非人类对齐。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.29141 2026-05-29 cs.IR cs.AI 79%

Toward User Preference Alignment in LLM Recommendation via Explicit Context Feedback

通过显式上下文反馈实现LLM推荐中的用户偏好对齐

Weizhi Zhang, Wooseong Yang, Yuxin Cui, Zhaohui Guo, Hins Hu, Liangwei Yang, Henry Peng Zou, Qifei Wang, Hanqing Zeng, Jiayi Liu, Yinglong Xia, Philip S. Yu

机构 * University of Illinois Chicago(伊利诺伊大学芝加哥分校) Meta

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.AI

AI总结 本文主张在基于大语言模型的推荐系统中优先利用显式上下文反馈(如评论文本)来对齐用户偏好,提升推荐的个性化和可解释性。

Comments Published in CogMI 2025. https://ieeexplore.ieee.org/abstract/document/11417068

详情

展开后加载摘要…

URL PDF HTML 收藏