arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-10-27 至 2025-10-27 共收录 5 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 偏好对齐 5 篇

2505.17859 2025-10-27 cs.LG cs.AI stat.ML 86%

Scalable Valuation of Human Feedback through Provably Robust Model Alignment

Masahiro Fujisawa, Masaki Adachi, Michael A. Osborne

机构 * The University of Osaka(大阪大学) Lattice Lab, Toyota Motor Corporation(丰田公司Lattice实验室) Machine Learning Research Group, University of Oxford(牛津大学机器学习研究组) RIKEN AIP(理化学研究所AIP)

专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);DPO(abstract);分类 cs.AI、cs.LG

Comments Accepted by the 39th Conference on Neural Information Processing Systems (NeurIPS2025), 49 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21798 2025-10-27 cs.CL cs.AI 81%

Evaluating and Improving Cultural Awareness of Reward Models for LLM Alignment

Hongbin Zhang, Kehai Chen, Xuefeng Bai, Yang Xiang, Min Zhang

机构 * Institute of Computing and Intelligence, Harbin Institute of Technology, Shenzhen, China(计算与智能研究所,哈尔滨工业大学,深圳,中国) Peng Cheng Laboratory, Shenzhen, China(鹏城实验室,深圳,中国)

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.CL、cs.AI

Comments Under review;Work in progress;

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.12854 2025-10-27 cs.CL cs.AI 73%

TPO: Aligning Large Language Models with Multi-branch & Multi-step Preference Trees

Weibin Liao, Xu Chu, Yasha Wang

机构 * School of Computer Science, Peking University(北京大学计算机科学系) Center on Frontiers of Computing Studies, Peking University(北京大学前沿计算研究中心) National Research and Engineering Center of Software Engineering, Peking University(北京大学软件工程研究中心)

专题命中 偏好对齐 :alignment(abstract);DPO(abstract);分类 cs.CL、cs.AI

Comments Accepted by ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.11475 2025-10-27 cs.CL cs.AI cs.LG 67%

HelpSteer3-Preference: Open Human-Annotated Preference Data across Diverse Tasks and Languages

Zhilin Wang, Jiaqi Zeng, Olivier Delalleau, Hoo-Chang Shin, Felipe Soares, Alexander Bukharin, Ellie Evans, Yi Dong, Oleksii Kuchaiev

机构 * NVIDIA

专题命中 偏好对齐 :RLHF(abstract);分类 cs.CL、cs.AI、cs.LG

Comments NeurIPS 2025 Datasets and Benchmarks Track Camera Ready, 46 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.11080 2025-10-27 cs.CL cs.AI cs.LG 67%

BLEUBERI: BLEU is a surprisingly effective reward for instruction following

Yapei Chang, Yekyung Kim, Michael Krumdick, Amir Zadeh, Chuan Li, Chris Tanner, Mohit Iyyer

机构 * University of Maryland, College Park(马里兰大学学院公园分校)

专题命中 偏好对齐 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments neurips cam-ready

详情

展开后加载摘要…

URL PDF HTML 收藏