arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-09-16 至 2025-09-16 共收录 7 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全训练 7 篇

2504.16980 2025-09-16 cs.LG 83%

Safety Pretraining: Toward the Next Generation of Safe AI

Pratyush Maini, Sachin Goyal, Dylan Sam, Alex Robey, Yash Savani, Yiding Jiang, Andy Zou, Matt Fredrikson, Zacharcy C. Lipton, J. Zico Kolter

专题命中 安全训练 :safety(title,abstract);alignment(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.08486 2025-09-16 cs.CL 77%

Too Helpful, Too Harmless, Too Honest or Just Right?

Gautam Siddharth Kashyap, Mark Dras, Usman Naseem

机构 * School of Computing, Macquarie University(计算机学院,麦考瑞大学)

专题命中 安全训练 :alignment(abstract);safety(abstract);harmlessness(abstract);分类 cs.CL

Comments EMNLP'25 Main

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11431 2025-09-16 cs.AI cs.CL 62%

Securing AI Agents: Implementing Role-Based Access Control for Industrial Applications

Aadil Gani Ganie

专题命中 安全训练 :prompt injection(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.02465 2025-09-16 cs.CL cs.AI 62%

Revealing the Inherent Instructability of Pre-Trained Language Models

Seokhyun An, Minji Kim, Hyounghun Kim

机构 * Department of Computer Science and Engineering, UNIST(UNIST计算机科学与工程系) Graduate School of Artificial Intelligence, POSTECH(POSTECH人工智能研究生院) Department of Computer Science and Engineering, POSTECH(POSTECH计算机科学与工程系)

专题命中 安全训练 :safety(abstract);分类 cs.CL、cs.AI

Comments Findings of EMNLP 2025 (32 pages). Code available at https://github.com/seokhyunan/response-tuning

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15776 2025-09-16 cs.CL cs.IR 57%

ConvSearch-R1: Enhancing Query Reformulation for Conversational Search with Reasoning via Reinforcement Learning

Changtai Zhu, Siyin Wang, Ruijun Feng, Kai Song, Xipeng Qiu

机构 * Fudan University(复旦大学) ByteDance Inc(字节跳动公司) University of New South Wales(新南威尔士大学)

专题命中 安全训练 :alignment(abstract);分类 cs.CL

Comments Accepted by EMNLP 2025 at the Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10478 2025-09-16 cs.NI cs.LG cs.SY eess.SY 57%

The LLM as a Network Operator: A Vision for Generative AI in the 6G Radio Access Network

Oluwaseyi Giwa, Michael Adewole, Tobi Awodumila, Pelumi Aderinto

机构 * African Institute for Mathematical Sciences(非洲数学科学研究所)

专题命中 安全训练 :safety(abstract);分类 cs.LG

Comments Submitted to Workshop on AI and ML for Next-Generation Wireless Communications and Networking, NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12085 2025-09-16 eess.SY cs.SY 50%

Compositional shield synthesis for safe reinforcement learning in partial observability

Steven Carr, Georgios Bakirtzis, Ufuk Topcu

专题命中 安全训练 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏