arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-09-18 至 2025-09-18 共收录 5 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 偏好对齐 5 篇

2405.13541 2025-09-18 cs.CL cs.AI cs.LG 78%

Annotation-Efficient Language Model Alignment via Diverse and Representative Response Texts

Yuu Jinnai, Ukyo Honda

机构 * CyberAgent / Tokyo, Japan(CyberAgent)

专题命中 偏好对齐 :alignment(title);分类 cs.CL、cs.AI、cs.LG

Comments EMNLP Findings, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15495 2025-09-18 cs.SE 67%

SynthCoder: A Synthetical Strategy to Tune LLMs for Code Completion

Dongjun Yu, Xiao Yan, Zhenrui Li, Jipeng Xiao, Haochuan He, Yongda Yu, Hao Zhang, Guoping Rong, Xiaobo Huang

专题命中 偏好对齐 :alignment(abstract);DPO(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.01658 2025-09-18 cs.LG cs.AI cs.IR 62%

CoPL: Collaborative Preference Learning for Personalizing LLMs

Youngbin Choi, Seunghyuk Cho, Minjong Lee, MoonJeong Park, Yesong Ko, Jungseul Ok, Dongwoo Kim

机构 * Graduate School of Artificial Intelligence, POSTECH(POSTECH人工智能研究生院) Department of Computer Science and Engineering, POSTECH(POSTECH计算机科学与工程系)

专题命中 偏好对齐 :alignment(abstract);分类 cs.AI、cs.LG

Comments 19pages, 13 figures, 11 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.13869 2025-09-18 cs.CL 57%

Do LLMs Align Human Values Regarding Social Biases? Judging and Explaining Social Biases with LLMs

Yang Liu, Chenhui Chu

专题命中 偏好对齐 :alignment(abstract);分类 cs.CL

Comments 38 pages, 31 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.06652 2025-09-18 cs.CL 57%

IntrEx: A Dataset for Modeling Engagement in Educational Conversations

Xingwei Tan, Mahathi Parvatham, Chiara Gambi, Gabriele Pergola

机构 * Department of Computer Science, University of Warwick, UK(沃里克大学计算机科学系) School of Computer Science, University of Sheffield, UK(谢菲尔德大学计算机科学学院) Department of Psychology, University of Warwick, UK(沃里克大学心理学系)

专题命中 偏好对齐 :RLHF(abstract);分类 cs.CL

Comments EMNLP 2025 Findings camera-ready, 9+7 pages

详情

展开后加载摘要…

URL PDF HTML 收藏