arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-10-30 至 2025-10-30 共收录 6 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 偏好对齐 6 篇

2510.23965 2025-10-30 cs.AI cs.LG stat.ML 84%

The Sign Estimator: LLM Alignment in the Face of Choice Heterogeneity

Ali Aouad, Aymane El Gadarri, Vivek F. Farias

机构 * MIT(麻省理工学院)

专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01183 2025-10-30 cs.LG cs.AI stat.ML 81%

Doubly Robust Alignment for Large Language Models

Erhan Xu, Kai Ye, Hongyi Zhou, Luhan Zhu, Francesco Quinzan, Chengchun Shi

机构 * Department of Statistics(统计系) LSE London, UK(伦敦大学学院) Department of Mathematics(数学系) Tsinghua University(清华大学) School of Design LCC, UAL London, UK(伦敦艺术大学设计学院) Department of Engineering Science(工程科学系) University of Oxford(牛津大学)

专题命中 偏好对齐 :alignment(title);RLHF(abstract);分类 cs.AI、cs.LG

Comments Accepted to NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.03690 2025-10-30 cs.CL 77%

Robust Preference Optimization via Dynamic Target Margins

Jie Sun, Junkang Wu, Jiancan Wu, Zhibo Zhu, Xingyu Lu, Jun Zhou, Lintao Ma, Xiang Wang

专题命中 偏好对齐 :alignment(abstract);DPO(abstract);safety(abstract);分类 cs.CL

Comments 18 pages, 6 figures, accepted to Findings of the 63rd Annual Meeting of the Association for Computational Linguistics (ACL 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.02745 2025-10-30 cs.AI cs.CL 76%

CURATRON: Complete and Robust Preference Data for Rigorous Alignment of Large Language Models

Son The Nguyen, Niranjan Uma Naresh, Theja Tulabandhula

机构 * University of Illinois Chicago(伊利诺伊大学芝加哥分校) Independent Researcher(独立研究者)

专题命中 偏好对齐 :alignment(title);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.01308 2025-10-30 cs.AI cs.CL cs.DB 62%

GradeSQL: Test-Time Inference with Outcome Reward Models for Text-to-SQL Generation from Large Language Models

Mattia Tritto, Giuseppe Farano, Dario Di Palma, Gaetano Rossiello, Fedelucio Narducci, Dharmashankar Subramanian, Tommaso Di Noia

专题命中 偏好对齐 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.17220 2025-10-30 cs.CL 61%

RLAIF-V: Open-Source AI Feedback Leads to Super GPT-4V Trustworthiness

Tianyu Yu, Haoye Zhang, Qiming Li, Qixin Xu, Yuan Yao, Da Chen, Xiaoman Lu, Ganqu Cui, Yunkai Dang, Taiwen He, Xiaocheng Feng, Jun Song, Bo Zheng, Zhiyuan Liu, Tat-Seng Chua, Maosong Sun

机构 * Tsinghua University(清华大学) Shanghai Qi Zhi Institute(上海启智研究院) Harbin Institute of Technology(哈尔滨工业大学) Taobao & Tmall Group of Alibaba(阿里巴巴淘宝与天猫集团) Peng Cheng Laboratory(鹏城实验室) National University of Singapore(新加坡国立大学)

专题命中 偏好对齐 :alignment(abstract);分类 cs.CL;RLHF(comments)

Comments Project Website: https://github.com/RLHF-V/RLAIF-V

详情

展开后加载摘要…

URL PDF HTML 收藏