arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-11-20 至 2025-11-20 共收录 5 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 偏好对齐 5 篇

2507.20964 2025-11-20 cs.AI cs.CC cs.GT cs.LG cs.MA 84%

Core Safety Values for Provably Corrigible Agents

Aran Nayebi

机构 * Aran Nayebi(独立研究者)

专题命中 偏好对齐 :safety(title,abstract);RLHF(abstract);分类 cs.AI、cs.LG

Comments 14 pages. To appear in AAAI 2026 Machine Ethics Workshop (W37) Proceedings

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17184 2025-11-20 cs.CL 79%

Towards Alignment-Centric Paradigm: A Survey of Instruction Tuning in Large Language Models

Xudong Han, Junjie Yang, Tianyang Wang, Ziqian Bi, Xinyuan Song, Junfeng Hao, Junhao Song

机构 * Department of Informatics, University of Sussex(信息学院,苏塞克斯大学) Pingtan Research Institute, Xiamen University(平潭研究院,厦门大学) Department of Computer Science, University of Liverpool(计算机科学系,利物浦大学) Department of Computer Science, Purdue University(计算机科学系,普渡大学) Department of Computer Science, Emory University(计算机科学系,埃默里大学) AI Agent Lab, Vokram Group(AI代理实验室,Vokram集团) Department of Computing, Imperial College London(计算系,帝国理工学院伦敦分校)

专题命中 偏好对齐 :alignment(title);safety(abstract);分类 cs.CL

Comments 24 pages, 7 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.03527 2025-11-20 cs.CL cs.LG 62%

FANAL -- Financial Activity News Alerting Language Modeling Framework

Urjitkumar Patel, Fang-Chun Yeh, Chinmay Gondhalekar, Hari Nalluri

专题命中 偏好对齐 :alignment(abstract);分类 cs.CL、cs.LG

Comments Accepted for the IEEE International Workshop on Large Language Models for Finance, 2024. This is a preprint version

Journal ref 2024 IEEE International Conference on Big Data, Dec 15-18, 2024, Electronic ISBN: 979-8-3503-6248-0, Electronic ISSN: 2573-2978

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15573 2025-11-20 cs.CY 57%

Two-Faced Social Agents: Context Collapse in Role-Conditioned Large Language Models

Vikram K Suresh

专题命中 偏好对齐 :alignment(abstract);分类 cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15038 2025-11-20 cs.SD cs.AI eess.AS 57%

Aligning Generative Music AI with Human Preferences: Methods and Challenges

Dorien Herremans, Abhinaba Roy

专题命中 偏好对齐 :alignment(abstract);分类 cs.AI

Comments Accepted at the AAAI-2026 Senior Member Track

详情

展开后加载摘要…

URL PDF HTML 收藏