arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 3266 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 偏好对齐 3266 篇

2603.26680 2026-05-12 cs.CL cs.AI 76%

AlpsBench: An LLM Personalization Benchmark for Real-Dialogue Memorization and Preference Alignment

AlpsBench: 一个面向真实对话记忆与偏好对齐的LLM个性化基准

Jianfei Xiao, Xiang Yu, Chengbing Wang, Wuqiang Zheng, Xinyu Lin, Kaining Liu, Hongxun Ding, Yang Zhang, Wenjie Wang, Fuli Feng, Xiangnan He

机构 * University of Science and Technology of China(科学技术大学) National University of Singapore(新加坡国立大学)

专题命中 偏好对齐 :alignment(title);分类 cs.CL、cs.AI

AI总结 AlpsBench通过真实人类与LLM对话数据构建,包含2500个长期交互序列及验证的记忆,评估个性化信息提取、更新、检索与利用等核心任务,揭示LLM在记忆管理中的局限性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.05967 2026-05-12 cs.AI cs.LG stat.ML 76%

Preference Learning for AI Alignment: a Causal Perspective

偏好学习用于AI对齐:一种因果视角

Katarzyna Kobalczyk, Mihaela van der Schaar

机构 * Department of Applied Mathematics and Theoretical Physics(应用数学与理论物理系)

专题命中 偏好对齐 :alignment(title);分类 cs.AI、cs.LG

AI总结 本文从因果视角出发,探讨了基于偏好数据的奖励建模在对齐大语言模型与人类价值观中的关键作用,提出因果工具箱以解决因果误识别、偏好异质性和用户特定因素的混淆问题,并提出未来研究的方向。

Journal ref Proceedings of the 42nd International Conference on Machine Learning, Vancouver, Canada. PMLR 267, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18600 2026-04-13 cs.CV cs.AI cs.LG 76%

Chain-of-Zoom: Extreme Super-Resolution via Scale Autoregression and Preference Alignment

链式放大:通过尺度自回归和偏好对齐实现极端超分辨率

Bryan Sangwoo Kim, Jeongsol Kim, Jong Chul Ye

机构 * KAIST AI(韩国科学技术院人工智能系)

专题命中 偏好对齐 :alignment(title);分类 cs.AI、cs.LG

AI总结 本文提出Chain-of-Zoom框架,通过自回归链和多尺度提示实现极端超分辨率,利用GRPO优化文本提示对齐人类偏好,实验显示标准4x扩散模型在CoZ下可实现256倍以上高质量放大。

Comments NeurIPS 2025 (Spotlight)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14628 2026-04-08 cs.CL cs.AI 76%

RLAIF-SPA: Structured AI Feedback for Semantic-Prosodic Alignment in Speech Synthesis

RLAIF-SPA: 结构化AI反馈用于语音合成中的语义-语调对齐

Qing Yang, Zhenghao Liu, Yangfan Du, Pengcheng Huang, Tong Xiao

机构 * School of Computer Science and Engineering, Northeastern University, China(东北大学计算机科学与工程学院)

专题命中 偏好对齐 :alignment(title);分类 cs.CL、cs.AI

AI总结 本文提出RLAIF-SPA框架,通过整合强化学习从AI反馈来优化语音合成中的情感表达和可懂度,实验显示其在多个数据集上均取得显著提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.08777 2026-03-30 cs.CL cs.AI 76%

Fluent Alignment with Disfluent Judges: Post-training for Lower-resource Languages

流畅对齐与不流畅评判:低资源语言的后训练方法

David Samuel, Lilja Øvrelid, Erik Velldal, Andrey Kutuzov

机构 * Language Technology Group, University of Oslo(奥斯陆大学语言技术组)

专题命中 偏好对齐 :alignment(title);分类 cs.CL、cs.AI

AI总结 本文提出一种低资源语言的后训练方法,通过不流畅评判模型保持语言模型的流畅性,通过挪威语案例研究验证了基于策略的训练方法在无需额外数据时的有效性。

Journal ref The Fourteenth International Conference on Learning Representations (ICLR 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.09802 2025-11-04 cs.CL cs.AI 76%

Enhancing Reasoning Abilities of Small LLMs with Cognitive Alignment

Wenrui Cai, Chengyu Wang, Junbing Yan, Jun Huang, Xiangzhong Fang

机构 * Shanghai Jiao Tong University(上海交通大学) Alibaba Cloud Computing(阿里云计算)

专题命中 偏好对齐 :alignment(title);分类 cs.CL、cs.AI

Comments emnlp 2025 main conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.02745 2025-10-30 cs.AI cs.CL 76%

CURATRON: Complete and Robust Preference Data for Rigorous Alignment of Large Language Models

Son The Nguyen, Niranjan Uma Naresh, Theja Tulabandhula

机构 * University of Illinois Chicago(伊利诺伊大学芝加哥分校) Independent Researcher(独立研究者)

专题命中 偏好对齐 :alignment(title);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14400 2025-10-21 cs.CL cs.AI cs.IR 76%

MedTrust-RAG: Evidence Verification and Trust Alignment for Biomedical Question Answering

Yingpeng Ning, Yuanyuan Sun, Ling Luo, Yanhua Wang, Yuchen Pan, Hongfei Lin

机构 * College of Computer Science and Technology, Dalian University of Technology(大连理工大学计算机科学与技术学院) Air Force Communications NCO Academy(空军通信NCO学院)

专题命中 偏好对齐 :alignment(title);分类 cs.CL、cs.AI

Comments Accepted as a short paper at BlBM2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05283 2025-10-08 cs.AI cs.CL cs.CV 76%

Beyond Monolithic Rewards: A Hybrid and Multi-Aspect Reward Optimization for MLLM Alignment

Radha Gulhane, Sathish Reddy Indurthi

机构 * Radha Gulhane(独立研究者) Sathish Reddy Indurthi(独立研究者)

专题命中 偏好对齐 :alignment(title);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25568 2025-10-01 cs.CL cs.AI 76%

Probing the Limits of Stylistic Alignment in Vision-Language Models

Asma Farajidizaji, Akash Gupta, Vatsal Raina

机构 * Imperial College London(伦敦帝国学院) Apta AI

专题命中 偏好对齐 :alignment(title);分类 cs.CL、cs.AI

Comments 5 pages, 1 figure, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.11870 2025-07-15 cs.CL cs.AI 76%

Intuitive Fine-Tuning: Towards Simplifying Alignment into a Single Process

Ermo Hua, Biqing Qi, Kaiyan Zhang, Kai Tian, Xingtai Lv, Ning Ding, Bowen Zhou

机构 * Department of Electronic Engineering, Tsinghua University(清华大学电子工程系) Shanghai AI Laboratory(上海人工智能实验室)

专题命中 偏好对齐 :alignment(title);分类 cs.CL、cs.AI

Comments Accepted to ACL 2025, Oral & Panel Discussion

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23848 2025-06-02 cs.CL cs.LG 76%

Derailing Non-Answers via Logit Suppression at Output Subspace Boundaries in RLHF-Aligned Language Models

Harvey Dam, Jonas Knochelmann, Vinu Joseph, Ganesh Gopalakrishnan

专题命中 偏好对齐 :RLHF(title);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.10460 2025-05-29 cs.CL cs.LG 76%

Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond

Liang Wen, Yunke Cai, Fenrui Xiao, Xin He, Qi An, Zhenyu Duan, Yimin Du, Junchen Liu, Lifu Tang, Xiaowei Lv, Haosheng Zou, Yongchao Deng, Shousheng Jia, Xiangzheng Zhang

机构 * First Author Affiliation(第一作者机构)

专题命中 偏好对齐 :DPO(title);分类 cs.CL、cs.LG

Comments v4: ACL'25 industry track camera ready; v3: minor modifications; v2: better writing & format for later submission; all release at https://github.com/Qihoo360/Light-R1

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.11410 2024-10-16 cs.CL cs.AI 76%

PMMT: Preference Alignment in Multilingual Machine Translation via LLM Distillation

Shuqiao Sun, Yutong Yao, Peiwen Wu, Feijun Jiang, Kaifu Zhang

专题命中 偏好对齐 :alignment(title);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.17760 2024-01-03 cs.CL cs.LG 76%

Language Models are Bounded Pragmatic Speakers: Understanding RLHF from a Bayesian Cognitive Modeling Perspective

Khanh Nguyen

专题命中 偏好对齐 :RLHF(title);分类 cs.CL、cs.LG

Comments Proceedings of the First Workshop on Theory of Mind in Communicating Agents at (TOM @ ICML 2023)

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.04373 2023-10-11 cs.LG cs.AI 76%

Confronting Reward Model Overoptimization with Constrained RLHF

Ted Moskovitz, Aaditya K. Singh, DJ Strouse, Tuomas Sandholm, Ruslan Salakhutdinov, Anca D. Dragan, Stephen McAleer

专题命中 偏好对齐 :RLHF(title);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.14375 2022-09-30 cs.LG cs.CL 76%

Improving alignment of dialogue agents via targeted human judgements

Amelia Glaese, Nat McAleese, Maja Trębacz, John Aslanides, Vlad Firoiu, Timo Ewalds, Maribeth Rauh, Laura Weidinger, Martin Chadwick, Phoebe Thacker, Lucy Campbell-Gillingham, Jonathan Uesato, Po-Sen Huang, Ramona Comanescu, Fan Yang, Abigail See, Sumanth Dathathri, Rory Greig, Charlie Chen, Doug Fritz, Jaume Sanchez Elias, Richard Green, Soňa Mokrá, Nicholas Fernando, Boxi Wu, Rachel Foley, Susannah Young, Iason Gabriel, William Isaac, John Mellor, Demis Hassabis, Koray Kavukcuoglu, Lisa Anne Hendricks, Geoffrey Irving

专题命中 偏好对齐 :alignment(title);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.04910 2025-03-10 cs.CL stat.ME 76%

Maximizing Signal in Human-Model Preference Alignment

Kelsey Kraus, Margaret Kroll

专题命中 偏好对齐 :alignment(title,comments);分类 cs.CL

Comments Presented at AAAI 2025, special track on AI Alignment

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.05882 2026-08-25 cs.CL cs.AI cs.LG 版本更新 75%

An Empirical Study on Preference Tuning Generalization and Diversity Under Domain Shift

关于领域转移下偏好微调泛化性和多样性的实证研究

Constantinos Karouzos, Xingwei Tan, Nikolaos Aletras

机构 * School of Computer Science University of Sheffield, UK(计算机科学学院 谢菲尔德大学)

专题命中 偏好对齐 :alignment(abstract);safety(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本研究通过比较不同对齐目标和适应策略,探讨领域转移下偏好微调的泛化性和多样性,发现伪标签法能有效缓解领域转移带来的性能下降。

Comments Accepted to EMNLP 2026 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.06337 2026-08-06 cs.CL cs.AI cs.LG 版本更新 75%

Can Post-Training Transform LLMs into Causal Reasoners?

训练后能否将大语言模型转变为因果推理者?

Junqi Chen, Sirui Chen, Chaochao Lu

机构 * Shanghai Artificial Intelligence Library(上海人工智能图书馆) Fudan University(复旦大学) Tongji University(同济大学)

专题命中 偏好对齐 :DPO(abstract,abstract_cn);分类 cs.CL、cs.AI、cs.LG

AI总结 本文通过CauGym数据集评估了训练后方法对LLM因果推理能力的提升,发现适当训练可使小型模型在因果推理任务中超越大模型。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.00419 2026-08-04 cs.LG cs.AI cs.CL cs.IR 新提交 75%

Unleashing the Potential of Large Language Models: A Blueprint for Real-Time, Enterprise-Ready Deployments

释放大语言模型的潜力:面向实时、企业级部署的蓝图

Muhammad Faizan Raza, Shuo, Yang, Satish Mahadevan Srinivasan, Joanna F. DeFranco

专题命中 偏好对齐 :RLHF(abstract,abstract_cn);分类 cs.CL、cs.AI、cs.LG

AI总结 针对实时受监管场景中大语言模型的问题,提出基于模式的LLMOps架构,整合多模块并实现四项模式化贡献,优化权衡同时支持高风险领域的可审计与回滚部署。

Comments 6 pages, 1 figure. Authors' accepted version of an article published in IEEE Computer. The version of record is available at the DOI below

Journal ref Computer, vol. 59, no. 4, pp. 195-199, April 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.25091 2026-07-29 cs.AI cs.CL cs.LG math.OC stat.CO 新提交 75%

Towards Robust Reinforcement Learning for Small-Scale Language Model Agents

迈向用于小规模语言模型智能体的稳健强化学习

Md Rezwanul Haque, Md. Milon Islam, Fakhri Karray

机构 * University of Waterloo(滑铁卢大学) Khulna University of Engineering & Technology(库尔纳工程技术大学) Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)

专题命中 偏好对齐 :alignment(abstract);safety(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 研究小规模语言模型强化学习不稳定问题,识别出三种失败模式,提出容量余量假设,采用合并并重新初始化适配器技术等方法,所提系统稳定收敛,提升偏好胜率,优于指令调整基线且减少训练数据。

Comments Proceedings of the 2026 IEEE International Conference on Systems, Man, and Cybernetics (SMC), Bellevue, WA, USA

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.06393 2026-07-21 cs.AI cs.CL cs.LG cs.LO 75%

Conflict-Aware Fusion: Mitigating Logic Inertia in Large Language Models via Structured Cognitive Priors

冲突感知融合:通过结构化认知先验缓解大语言模型中的逻辑惯性

Qiming Bao, Xiaoxuan Fu, Michael Witbrock

机构 * Xtracta & Strong AI Lab, University of Auckland(Xtracta与强人工智能实验室,奥克兰大学) School of Humanities, China University of Political Science and Law(人文学院,中国政法大学) Strong AI Lab, University of Auckland(强人工智能实验室,奥克兰大学)

专题命中 偏好对齐 :DPO(abstract,abstract_cn);分类 cs.CL、cs.AI、cs.LG

AI总结 针对大语言模型在规则系统结构扰动下表现脆弱的问题,提出冲突感知融合训练流程,通过验证-演绎结构先验和符号推理奖励,在多个压力测试中实现鲁棒性饱和。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.11948 2026-07-15 cs.AI cs.CL cs.LG cs.MA 新提交 75%

Ontology-Amplified Distillation and Contextuality Auditing for Sovereign Enterprise Language Models: A Combined Proof-of-Mechanism and Negative-Results Method Study

主权企业语言模型的本体增强蒸馏与上下文审查:机制验证与负面结果的组合方法研究

Thanh Luong Tuan

机构 * AgenticOS(智能操作系统)

专题命中 偏好对齐 :DPO(abstract);safety(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 研究针对受数据驻留规则约束的金融机构需求,结合本体增强蒸馏机制验证与上下文审查方法,对Qwen3.6 - 27B学生模型进行训练及测试,结果不支持模型在多方面的优势,为企业语言模型应用提供参考。

Comments 15 pages, 2 figures. Combined proof-of-mechanism and negative-results method article consolidating ontology-amplified distillation with contextuality-audit routing for enterprise agents

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.18037 2026-07-07 cs.LG cs.AI cs.CL 版本更新 75%

Gradient Regularization Mitigates Reward Hacking in Reinforcement Learning from Human Feedback and Verifiable Rewards

梯度正则化减轻基于人类反馈和可验证奖励的强化学习中的奖励作弊

Johannes Ackermann, Michael Noukhovitch, Takashi Ishida, Masashi Sugiyama

机构 * The University of Tokyo(东京大学) RIKEN AIP(理化学研究所AIP) Mila(蒙特利尔大学Mila)

专题命中 偏好对齐 :RLHF(abstract,abstract_cn);分类 cs.CL、cs.AI、cs.LG

AI总结 研究基于人类反馈或可验证奖励的强化学习中奖励作弊问题,提出用梯度正则化使训练偏向奖励更准确区域,理论推导并实证验证,还改进方法,结果显示其比KL惩罚表现更好。

Comments Accepted at ICML 2026, 25 pages, 15 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.00539 2026-06-30 cs.LG cs.AI cs.CL 75%

Distributionally Robust Reinforcement Learning with Human Feedback

具有人类反馈的分布鲁棒强化学习

Debmalya Mandal, Paulius Sasnauskas, Goran Radanovic

专题命中 偏好对齐 :RLHF(abstract);DPO(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文提出一种分布鲁棒的强化学习方法,通过改进奖励基强化学习和直接偏好优化,提升模型在分布变化时的鲁棒性,实验显示在非分布任务中性能显著提升。

Comments Accepted at ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.26269 2026-06-05 cs.CL cs.AI cs.LG 75%

Calibrated Surprise: An Information-Theoretic Account of Creative Quality

校准的惊喜:一种信息论视角下的创造性质量

Bo Zou, Chao Xu

机构 * Bo Zou(邹波) Chao Xu(徐超)

专题命中 偏好对齐 :RLHF(abstract,abstract_cn);分类 cs.CL、cs.AI、cs.LG

AI总结 本文提出了一种信息论框架,用于评估创造性写作的质量,通过校准的惊喜概念,结合香农互信息理论,量化了高质量文本与降质文本之间的差异。

Comments 28 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.04284 2026-06-04 cs.LG cs.AI cs.CL 75%

Sparse Mixture-of-Experts Reward Models Learn Interpretable and Specialized Experts for Personalized Preference Modeling

稀疏混合专家奖励模型学习可解释且专业化的专家用于个性化偏好建模

Yifan Wang, Jinyi Mu, Mayank Jobanputra, Yu Wang, Ji-Ung Lee, Soyoung Oh, Isabel Valera, Vera Demberg

机构 * Saarland University(萨尔兰大学) Independent Researcher(独立研究者) Bielefeld University(比勒菲尔德大学) Max Planck Institute for Software Systems(马克斯·普朗克软件系统研究所) Max Planck Institute for Informatics(马克斯·普朗克信息研究所)

专题命中 偏好对齐 :RLHF(abstract,abstract_cn);分类 cs.CL、cs.AI、cs.LG

AI总结 提出稀疏混合专家奖励模型,通过稀疏路由和专家多样性训练,从二元偏好数据中学习可解释的专家模式,提升个性化偏好建模的测试时适应性和可解释性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.01811 2026-06-02 cs.CL cs.AI cs.LG 75%

"I've Seen How This Goes": Characterizing Diversity via Progressive Conditional Surprise

“我知道这会如何发展”:通过渐进条件惊奇度刻画多样性

Matthew Khoriaty, David Williams-King, Shi Feng

机构 * University of California, Berkeley(加州大学伯克利分校) University of Cambridge(剑桥大学) Stanford University(斯坦福大学)

专题命中 偏好对齐 :DPO(abstract,abstract_cn);分类 cs.CL、cs.AI、cs.LG

AI总结 提出一种基于上下文学习的多样性度量方法 Decan(D_{Ca_n}),通过单次前向传递计算每个字节的得分,无需嵌入模型、参考语料或人工标注,在多个基准上验证了其有效性。

Comments 28 pages, 18 figures, 9 tables. Accepted to the Workshop on Generative AI, Creativity, and Human-AI Co-Creation @ ICML 2026 (non-archival). Code and data: https://github.com/AMindToThink/icl-diversity

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01009 2026-06-02 cs.CV cs.MM 75%

POVQA: Preference-Optimized Video Question Answering with Rationales for Data Efficiency

POVQA: 基于偏好的视频问答与数据效率的推理

Ashim Dahal, Ankit Ghimire, Saydul Akbar Murad, Nick Rahimi

机构 * University of Southern Mississippi(密西根州立大学)

专题命中 偏好对齐 :DPO(abstract,abstract_cn);alignment(abstract)

AI总结 提出POVQA方法,通过时间池化压缩视频帧、监督微调加偏好优化,在长视频问答中实现数据高效推理。

Comments Accepted in MAR at CVPR Workshop (Proceedings Track)

Journal ref Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops, 2026, pp. 11533-11542

详情

展开后加载摘要…

URL PDF HTML 收藏