arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-09-19 至 2025-09-19 共收录 4 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 偏好对齐 4 篇

2507.05129 2025-09-19 cs.CL cs.CY cs.LG 75%

SMART: Simulated Students Aligned with Item Response Theory for Question Difficulty Prediction

Alexander Scarlatos, Nigel Fernandez, Christopher Ormerod, Susan Lottridge, Andrew Lan

机构 * University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校) Cambium Assessment(Cambium评估)

专题命中 偏好对齐 :alignment(abstract);DPO(abstract);分类 cs.CL、cs.CY、cs.LG

Comments Published in EMNLP 2025: The 2025 Conference on Empirical Methods in Natural Language Processing

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.20415 2025-09-19 cs.CL 57%

Enhancing Logical Reasoning in Language Models via Symbolically-Guided Monte Carlo Process Supervision

Xingwei Tan, Marco Valentino, Mahmud Akhter, Maria Liakata, Nikolaos Aletras

机构 * School of Computer Science, University of Sheffield(谢菲尔德大学计算机科学学院) School of Electronic Engineering and Computer Science, Queen Mary University of London(伦敦大学玛丽女王学院电子工程与计算机科学学院) The Alan Turing Institute(艾伦·图灵研究所)

专题命中 偏好对齐 :DPO(abstract);分类 cs.CL

Comments EMNLP 2025 (Main), 9+6 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19176 2025-09-19 cs.CL 57%

Assistant-Guided Mitigation of Teacher Preference Bias in LLM-as-a-Judge

Zhuo Liu, Moxin Li, Xun Deng, Qifan Wang, Fuli Feng

机构 * University of Science and Technology of China(中国科学技术大学) National University of Singapore(新加坡国立大学) Meta AI

专题命中 偏好对齐 :alignment(abstract);分类 cs.CL

Comments Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.05154 2025-09-19 cs.CL 57%

CARE: Multilingual Human Preference Learning for Cultural Awareness

Geyang Guo, Tarek Naous, Hiromi Wakaki, Yukiko Nishimura, Yuki Mitsufuji, Alan Ritter, Wei Xu

机构 * Georgia Institute of Technology(佐治亚理工学院) Sony Group Corporation(索尼集团) Sony AI(索尼人工智能)

专题命中 偏好对齐 :alignment(abstract);分类 cs.CL

Comments Accepted at EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏