arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-10-03 至 2025-10-03 共收录 10 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 偏好对齐 10 篇

2510.01700 2025-10-03 cs.AI cs.CV cs.LG 81%

VaPR -- Vision-language Preference alignment for Reasoning

Rohan Wadhawan, Fabrice Y Harel-Canada, Zi-Yi Dou, Suhaila Shakiah, Robinson Piramuthu, Nanyun Peng

机构 * Department of Computer Science, University of California Los Angeles(加州大学洛杉矶分校计算机科学系) Amazon.com, Inc.(亚马逊公司)

专题命中 偏好对齐 :alignment(title);DPO(abstract);分类 cs.AI、cs.LG

Journal ref COLM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01616 2025-10-03 cs.CL 80%

Efficient Training of Robust Traditional Chinese LLaMA-1B on a Single Consumer GPU: Continual Pre-training, SFT, and DPO

Yu-Cheng Chih, Ming-Tao Duan, Yong-Hao Hou

专题命中 偏好对齐 :DPO(title,abstract);分类 cs.CL

Comments 17 pages, 1 figures, 2 tables. Technical report. Introduces PureTC-1B, an adapter-based pipeline for stabilizing Small Language Models in Traditional Chinese using CPT, SFT, and DPO

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23761 2025-10-03 cs.LG cs.AI cs.CL 76%

Differential Information Distribution: A Bayesian Perspective on Direct Preference Optimization

Yunjae Won, Hyunji Lee, Hyeonbin Hwang, Minjoon Seo

机构 * KAIST AI(韩国科学技术院人工智能研究所)

专题命中 偏好对齐 :DPO(abstract,comments);alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Preprint, under review. 39 pages, 12 figures. Updates from v1: Added new theoretical results on DPO training dynamics and policy exploration, included experiments with Qwen3-4B, and refined the discussion of log-margin dynamics

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.03206 2025-10-03 cs.CL cs.AI 73%

Enhancing Personalized Multi-Turn Dialogue with Curiosity Reward

Yanming Wan, Jiaxing Wu, Marwa Abdulhai, Lior Shani, Natasha Jaques

机构 * Google DeepMind(谷歌DeepMind) University of Washington(华盛顿大学) Google Research(谷歌研究) University of California, Berkeley(伯克利加州大学)

专题命中 偏好对齐 :RLHF(abstract);safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.03323 2025-10-03 cs.CL cs.AI cs.LG 67%

Out-of-Distribution Detection using Synthetic Data Generation

Momin Abbas, Muneeza Azmat, Raya Horesh, Mikhail Yurochkin

专题命中 偏好对齐 :RLHF(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Accepted to COLM 2025. Camera-ready version

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01394 2025-10-03 cs.LG cs.CL 62%

Optimal Stopping vs Best-of-$N$ for Inference Time Optimization

Yusuf Kalayci, Vinod Raman, Shaddin Dughmi

机构 * University of Southern California(南加州大学) University of Michigan(密歇根大学)

专题命中 偏好对齐 :RLHF(abstract);分类 cs.CL、cs.LG

Comments 24 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01236 2025-10-03 cs.CL cs.LG 62%

GRPO++: Enhancing Dermatological Reasoning under Low Resource Settings

Ismam Nur Swapnil, Aranya Saha, Tanvir Ahmed Khan, Mohammad Ariful Haque

机构 * Department of Electrical and Electronic Engineering, Bangladesh University of Engineering and Technology (BUET)(电气与电子工程系,孟加拉国工程与技术大学)

专题命中 偏好对齐 :DPO(abstract);分类 cs.CL、cs.LG

Comments Will be submitted at IEEE JBHI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.03238 2025-10-03 cs.CL cs.AI 62%

FANS -- Formal Answer Selection for Natural Language Math Reasoning Using Lean4

Jiarui Yao, Ruida Wang, Tong Zhang

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 偏好对齐 :trustworthy(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22558 2025-10-03 cs.AI 57%

StepORLM: A Self-Evolving Framework With Generative Process Supervision For Operations Research Language Models

Chenyu Zhou, Tianyi Xu, Jianghao Lin, Dongdong Ge

机构 * Shanghai Jiao Tong University(上海交通大学)

专题命中 偏好对齐 :DPO(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.01051 2025-10-03 cs.CV 50%

Diffusion Model as a Noise-Aware Latent Reward Model for Step-Level Preference Optimization

Tao Zhang, Cheng Da, Kun Ding, Huan Yang, Kun Jin, Yan Li, Tingting Gao, Di Zhang, Shiming Xiang, Chunhong Pan

机构 * MAIS, CASIA(中国科学院自动化研究所MAIS实验室) Kuaishou Technology(快手科技) School of Artificial Intelligence, UCAS(中国科学院大学人工智能学院)

专题命中 偏好对齐 :alignment(abstract)

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏