Core Safety Values for Provably Corrigible Agents
机构 * Aran Nayebi(独立研究者)
专题命中 偏好对齐 :safety(title,abstract);RLHF(abstract);分类 cs.AI、cs.LG
Comments 14 pages. To appear in AAAI 2026 Machine Ethics Workshop (W37) Proceedings
AI 大模型
大模型对齐、安全、越狱、红队、提示注入和可信评测。
机构 * Aran Nayebi(独立研究者)
专题命中 偏好对齐 :safety(title,abstract);RLHF(abstract);分类 cs.AI、cs.LG
Comments 14 pages. To appear in AAAI 2026 Machine Ethics Workshop (W37) Proceedings
机构 * Department of Informatics, University of Sussex(信息学院,苏塞克斯大学) ; Pingtan Research Institute, Xiamen University(平潭研究院,厦门大学) ; Department of Computer Science, University of Liverpool(计算机科学系,利物浦大学) ; Department of Computer Science, Purdue University(计算机科学系,普渡大学) ; Department of Computer Science, Emory University(计算机科学系,埃默里大学) ; AI Agent Lab, Vokram Group(AI代理实验室,Vokram集团) ; Department of Computing, Imperial College London(计算系,帝国理工学院伦敦分校)
专题命中 偏好对齐 :alignment(title);safety(abstract);分类 cs.CL
Comments 24 pages, 7 figures, 5 tables
专题命中 偏好对齐 :alignment(abstract);分类 cs.CL、cs.LG
Comments Accepted for the IEEE International Workshop on Large Language Models for Finance, 2024. This is a preprint version
Journal ref 2024 IEEE International Conference on Big Data, Dec 15-18, 2024, Electronic ISBN: 979-8-3503-6248-0, Electronic ISSN: 2573-2978
专题命中 偏好对齐 :alignment(abstract);分类 cs.CY
专题命中 偏好对齐 :alignment(abstract);分类 cs.AI
Comments Accepted at the AAAI-2026 Senior Member Track