arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-11-18 至 2025-11-18 共收录 11 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 偏好对齐 11 篇

2511.09385 2025-11-18 cs.CL 83%

AMaPO: Adaptive Margin-attached Preference Optimization for Language Model Alignment

Ruibo Deng, Duanyu Feng, Wenqiang Lei

专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);分类 cs.CL

Comments AAAI 2026 AIA oral, our code is available at https://github.com/Shiroha-Offical/AMaPO

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13007 2025-11-18 cs.AI cs.LG 81%

GEM: Generative Entropy-Guided Preference Modeling for Few-shot Alignment of LLMs

Yiyang Zhao, Huiyu Bai, Xuejiao Zhao

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.AI、cs.LG

Comments This paper has been accepted by AAAI 2026-AIA and designated as an oral presentation paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12573 2025-11-18 cs.CL cs.AI 81%

Mitigating Length Bias in RLHF through a Causal Lens

Hyeonji Kim, Sujeong Oh, Sanghack Lee

专题命中 偏好对齐 :RLHF(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.06965 2025-11-18 cs.CL cs.AI 81%

Uncovering Factor Level Preferences to Improve Human-Model Alignment

Juhyun Oh, Eunsu Kim, Jiseon Kim, Wenda Xu, Inha Cha, William Yang Wang, Alice Oh

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09724 2025-11-18 cs.AI 79%

UDA: Unsupervised Debiasing Alignment for Pair-wise LLM-as-a-Judge

Yang Zhang, Cunxiang Wang, Lindong Wu, Wenbo Yu, Yidong Wang, Guangsheng Bao, Jie Tang

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13016 2025-11-18 cs.LG 70%

The Good, The Bad, and The Hybrid: A Reward Structure Showdown in Reasoning Models Training

Subramanyam Sahoo

机构 * Berkeley AI Safety Initiative (BASIS) UC Berkeley(伯克利人工智能安全计划(BASIS)加州大学伯克利分校)

专题命中 偏好对齐 :alignment(abstract);RLHF(abstract);分类 cs.LG

Comments Paper accepted to the 2nd Workshop on Aligning Reinforcement Learning Experimentalists and Theorists (ARLET 2025) at NeurIPS; the paper consists of 14 pages (including the appendix) and contains 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12036 2025-11-18 cs.CE cond-mat.mtrl-sci cs.AI cs.CL cs.LG 67%

Preference Learning from Physics-Based Feedback: Tuning Language Models to Design BCC/B2 Superalloys

Satanu Ghosh, Collin Holgate, Neal R. Brodnik, Doug Downey, Samantha Daly, Tresa M. Pollock, Samuel Carton

机构 * Department of Computer Science, University of New Hampshire(新罕布什尔大学计算机科学系) Materials Department, University of California, Santa Barbara(加州大学圣芭芭拉分校材料系) Department of Mechanical Engineering, University of California, Santa Barbara(加州大学圣芭芭拉分校机械工程系) Allen Institute for Artificial Intelligence(人工智能研究院)

专题命中 偏好对齐 :DPO(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12464 2025-11-18 cs.CL 57%

Probing Preference Representations: A Multi-Dimensional Evaluation and Analysis Method for Reward Models

Chenglong Wang, Yifu Huo, Yang Gan, Yongyu Mu, Qiaozhi He, Murun Yang, Bei Li, Chunliang Zhang, Tongran Liu, Anxiang Ma, Zhengtao Yu, Jingbo Zhu, Tong Xiao

机构 * School of Computer Science and Engineering, Northeastern University(东北大学计算机科学与工程学院) Meituan Inc.(美团公司) NiuTrans Research(牛译研所) CAS Key Laboratory of Behavioral Science, Institute of Psychology, CAS(中国科学院行为科学重点实验室) Kunming University of Science and Technology(昆明理工大学)

专题命中 偏好对齐 :alignment(abstract);分类 cs.CL

Comments Accepted by AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.10071 2025-11-18 cs.AI 57%

Advanced Tool Learning and Selection System (ATLASS): A Closed-Loop Framework Using LLM

Mohd Ariful Haque, Justin Williams, Sunzida Siddique, Md. Hujaifa Islam, Hasmot Ali, Kishor Datta Gupta, Roy George

机构 * Clark Atlanta University(克拉克亚特兰大大学) Daffodil International University(达福尔国际大学) Ahsanullah University of Science and Technology(阿沙努拉大学科学与技术学院)

专题命中 偏好对齐 :safety(abstract);分类 cs.AI

Journal ref 2025 IEEE International Conference on Service-Oriented System Engineering (SOSE), Tucson, AZ, USA, 21-24 July 2025, IEEE

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.18212 2025-11-18 cs.CL 57%

Better Language Model-Based Judging Reward Modeling through Scaling Comprehension Boundaries

Meiling Ning, Zhongbao Zhang, Junda Ye, Jiabao Guo, Qingyuan Guan

专题命中 偏好对齐 :RLHF(abstract);分类 cs.CL

Comments After further internal discussion, our author team has decided to withdraw this submission due to the need for several important refinements to the manuscript. All co-authors have been informed and agree with this decision

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07864 2025-11-18 cs.CV 50%

Tracing and Mitigating Hallucinations in Multimodal LLMs via Dynamic Attention Localization

Tiancheng Yang, Lin Zhang, Jiaye Lin, Guimin Hu, Di Wang, Lijie Hu

机构 * MBZUAI Provable Responsible AI and Data Analytics (PRADA) Lab(可证明负责任的人工智能与数据分析实验室) King Abdullah University of Science and Technology(卡迪夫大学科学与技术大学) School of Advanced Interdisciplinary Sciences, University of Chinese Academy of Sciences(中国科学院大学先进交叉科学学院) University of Copenhagen(哥本哈根大学) Tsinghua University(清华大学)

专题命中 偏好对齐 :DPO(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏