arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-11-11 至 2025-11-11 共收录 6 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全训练 6 篇

2501.13677 2025-11-11 cs.LG cs.CR 83%

HumorReject: Decoupling LLM Safety from Refusal Prefix via A Little Humor

Zihui Wu, Haichang Gao, Jiacheng Luo, Zhaoxiang Liu

专题命中 安全训练 :safety(title,abstract);alignment(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06262 2025-11-11 cs.AI cs.CY 62%

GAIA: A General Agency Interaction Architecture for LLM-Human B2B Negotiation & Screening

Siming Zhao, Qi Li

机构 * Alibaba.com US E-Commerce(阿里巴巴美国电子商务) Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06111 2025-11-11 cs.LG 57%

Guardian-regularized Safe Offline Reinforcement Learning for Smart Weaning of Mechanical Circulatory Devices

Aysin Tumay, Sophia Sun, Sonia Fereidooni, Aaron Dumas, Elise Jortberg, Rose Yu

机构 * University of California(加州大学) California Institute of Technology(加州理工学院) Abiomed(阿比omed)

专题命中 安全训练 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06094 2025-11-11 cs.LG 57%

Approximating Shapley Explanations in Reinforcement Learning

Daniel Beechey, Özgür Şimşek

机构 * University of Bath(巴斯大学)

专题命中 安全训练 :safety(abstract);分类 cs.LG

Comments Camera-ready version. Published at the Conference on Neural Information Processing Systems (NeurIPS 2025)

Journal ref Proceedings of the Conference on Neural Information Processing Systems (NeurIPS 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05532 2025-11-11 cs.CL 57%

Beyond One-Size-Fits-All: Personalized Harmful Content Detection with In-Context Learning

Rufan Zhang, Lin Zhang, Xianghang Mi

机构 * University of Science and Technology of China(科学技术大学)

专题命中 安全训练 :safety(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.01739 2025-11-11 cs.AI cs.LG 54%

Conceptual Belief-Informed Reinforcement Learning

Xingrui Gu, Chuyi Jiang, Laixi Shi

机构 * University of California, Berkeley(加州大学伯克利分校) Columbia University(哥伦比亚大学) Johns Hopkins University(约翰霍普金斯大学)

专题命中 安全训练 :alignment(comments,journal_ref);分类 cs.AI、cs.LG

Comments Accepted by ICML 2025 Workshop on Models of Human Feedback for AI Alignment

Journal ref ICML 2025 Workshop on Models of Human Feedback for AI Alignment

详情

展开后加载摘要…

URL PDF HTML 收藏