arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-11-04 至 2025-11-04 共收录 7 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全训练 7 篇

2511.01689 2025-11-04 cs.CL cs.AI cs.LG 67%

Open Character Training: Shaping the Persona of AI Assistants through Constitutional AI

Sharan Maiya, Henning Bartsch, Nathan Lambert, Evan Hubinger

专题命中 安全训练 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments 12 pages, 6 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00993 2025-11-04 cs.AI cs.LG 62%

Aligning LLM agents with human learning and adjustment behavior: a dual agent approach

Tianming Liu, Jirong Yang, Yafeng Yin, Manzi Li, Linghao Wang, Zheng Zhu

机构 * Department of Civil and Environmental Engineering, University of Michigan, Ann Arbor, United States(美国密歇根大学土木与环境工程系) Department of Electrical Engineering and Computer Science, University of Michigan, Ann Arbor, United States(美国密歇根大学电气工程与计算机科学系) College of Civil Engineering and Architecture, Zhejiang University, Hangzhou, China(浙江大学建筑工程学院)

专题命中 安全训练 :alignment(abstract);分类 cs.AI、cs.LG

Comments 32 pages, 6 figures, 7 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00880 2025-11-04 cs.LG cs.AI 62%

KFCPO: Kronecker-Factored Approximated Constrained Policy Optimization

Joonyoung Lim, Younghwan Yoo

机构 * School of Computer Science and Engineering, Pusan National University, Busan, Korea(计算机科学与工程学院,釜山国立大学,韩国釜山)

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG

Comments 12 pages, 8 figures, submitted to ECAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.16813 2025-11-04 cs.CL cs.AI 62%

Incivility and Rigidity: Evaluating the Risks of Fine-Tuning LLMs for Political Argumentation

Svetlana Churina, Kokil Jaidka

专题命中 安全训练 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00029 2025-11-04 cs.LG cs.AI 62%

Feature-Guided SAE Steering for Refusal-Rate Control using Contrasting Prompts

Samaksh Bhargav, Zining Zhu

机构 * Edison Academy Magnet School New Jersey, USA(新泽西州埃迪森磁校) Department of Computer Science(计算机科学系) Stevens Institute of Technology(史蒂文斯理工学院)

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG

Comments 12 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00112 2025-11-04 cs.RO cs.AI 57%

Real-DRL: Teach and Learn in Reality

Yanbing Mao, Yihao Cai, Lui Sha

机构 * Engineering Technology Division(工程科技部门) Wayne State University(韦恩州立大学) Department of Electrical and Computer Engineering(电气与计算机工程系) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 安全训练 :safety(abstract);分类 cs.AI

Comments 37 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16240 2025-11-04 cs.RO 50%

Cosmos-Surg-dVRK: World Foundation Model-based Automated Online Evaluation of Surgical Robot Policy Learning

Lukas Zbinden, Nigel Nelson, Juo-Tung Chen, Xinhao Chen, Ji Woong Kim, Mahdi Azizian, Axel Krieger, Sean Huver

机构 * NVIDIA Johns Hopkins University(约翰霍普金斯大学) Stanford University(斯坦福大学)

专题命中 安全训练 :alignment(abstract)

Comments minor metadata and notation fixes; +3 citations

详情

展开后加载摘要…

URL PDF HTML 收藏