arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-10-21 至 2025-10-21 共收录 11 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 偏好对齐 11 篇

2510.16167 2025-10-21 cs.LG cs.CL 86%

Alignment is Localized: A Causal Probe into Preference Layers

Archie Chaudhury

机构 * Independent(独立研究者)

专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);safety(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.01203 2025-10-21 cs.LG stat.ML 83%

KL-Regularized RLHF with Multiple Reference Models: Exact Solutions and Sample Complexity

Gholamali Aminian, Amir R. Asadi, Idan Shenfeld, Youssef Mroueh

机构 * The Alan Turing Institute(艾伦·图灵研究所) Statistical Laboratory(统计实验室) University of Cambridge(剑桥大学) Massachusetts Institute of Technology(麻省理工学院) IBM Research USA(IBM美国研究)

专题命中 偏好对齐 :RLHF(title,abstract);alignment(abstract);分类 cs.LG

Comments Extra experiments are added in new version

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.17041 2025-10-21 cs.CV cs.AI cs.LG 81%

Free$^2$Guide: Training-Free Text-to-Video Alignment using Image LVLM

Jaemin Kim, Bryan Sangwoo Kim, Jong Chul Ye

机构 * Graduate School of AI, KAIST(人工智能研究生院,韩国科学技术院)

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.AI、cs.LG

Comments ICCV 2025 accepted

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.05831 2025-10-21 cs.CL 79%

Leveraging Robust Optimization for LLM Alignment under Distribution Shifts

Mingye Zhu, Yi Liu, Zheren Fu, Yongdong Zhang, Zhendong Mao

机构 * University of Science and Technology of China(中国科学技术大学) State Key Laboratory of Communication Content Cognition(通信内容认知国家重点实验室)

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.CL

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.15065 2025-10-21 cs.LG 77%

Direct Preference Optimization With Unobserved Preference Heterogeneity: The Necessity of Ternary Preferences

Keertana Chidambaram, Karthik Vinay Seetharaman, Vasilis Syrgkanis

机构 * Stanford University(斯坦福大学)

专题命中 偏好对齐 :alignment(abstract);RLHF(abstract);DPO(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14400 2025-10-21 cs.CL cs.AI cs.IR 76%

MedTrust-RAG: Evidence Verification and Trust Alignment for Biomedical Question Answering

Yingpeng Ning, Yuanyuan Sun, Ling Luo, Yanhua Wang, Yuchen Pan, Hongfei Lin

机构 * College of Computer Science and Technology, Dalian University of Technology(大连理工大学计算机科学与技术学院) Air Force Communications NCO Academy(空军通信NCO学院)

专题命中 偏好对齐 :alignment(title);分类 cs.CL、cs.AI

Comments Accepted as a short paper at BlBM2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.20548 2025-10-21 cs.LG cs.AI cs.CL 75%

$Q\sharp$: Provably Optimal Distributional RL for LLM Post-Training

Jin Peng Zhou, Kaiwen Wang, Jonathan Chang, Zhaolin Gao, Nathan Kallus, Kilian Q. Weinberger, Kianté Brantley, Wen Sun

专题命中 偏好对齐 :alignment(abstract);DPO(abstract);分类 cs.CL、cs.AI、cs.LG

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12081 2025-10-21 cs.CV cs.AI cs.CL 73%

VimoRAG: Video-based Retrieval-augmented 3D Motion Generation for Motion Language Models

Haidong Xu, Guangwei Xu, Zhedong Zheng, Xiatian Zhu, Wei Ji, Xiangtai Li, Ruijie Guo, Meishan Zhang, Min zhang, Hao Fei

机构 * Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳)) University of Macau(澳门大学) University of Surrey(Surrey大学) Nanjing University(南京大学) Nanyang Technological University(南洋理工大学) National University of Singapore(新加坡国立大学)

专题命中 偏好对齐 :alignment(abstract);DPO(abstract);分类 cs.CL、cs.AI

Comments Accepted by NeurIPS 2025; Project Page: https://walkermitty.github.io/VimoRAG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12943 2025-10-21 cs.CL 57%

The Curious Case of Curiosity across Human Cultures and LLMs

Angana Borah, Zhijing Jin, Rada Mihalcea

机构 * University of Michigan - Ann Arbor(密歇根大学安阿伯分校) University of Toronto(多伦多大学) Vector Institute(向量研究所) MPI for Intelligent Systems, Tubingen, Germany(图宾根德国智能系统研究所)

专题命中 偏好对齐 :alignment(abstract);分类 cs.CL

Comments Preprint (Paper under review)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15993 2025-10-21 q-fin.PM cs.LG q-fin.ST 57%

Aligning Language Models with Investor and Market Behavior for Financial Recommendations

Fernando Spadea, Oshani Seneviratne

机构 * Rensselaer Polytechnic Institute(拉特兰理工学院)

专题命中 偏好对齐 :alignment(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.06023 2025-10-21 cs.CV 50%

Dual Caption Preference Optimization for Diffusion Models

Amir Saeidi, Yiran Luo, Agneet Chatterjee, Shamanthak Hegde, Bimsara Pathiraja, Yezhou Yang, Chitta Baral

机构 * School of Computing and Augmented Intelligence(计算与增强智能学院) Arizona State University(亚利桑那州立大学)

专题命中 偏好对齐 :DPO(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏