arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-08-18 至 2025-08-18 共收录 3 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全训练 3 篇

2411.02957 2025-08-18 cs.LG cs.SY eess.SY 79%

Embedding Safety into RL: A New Take on Trust Region Methods

Nikola Milosevic, Johannes Müller, Nico Scherf

机构 * Max Planck Institute for Human Cognitive and Brain Sciences(马克斯·普朗克人类认知与脑科学研究所) Center for Scalable Data Analytics and Artificial Intelligence(可扩展数据分析与人工智能中心)

专题命中 安全训练 :safety(title,abstract);分类 cs.LG

Comments Accepted at ICML 2025

Journal ref Proceedings of the 42nd International Conference on Machine Learning, Vancouver, Canada. PMLR 267, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09666 2025-08-18 cs.CL 70%

Slow Tuning and Low-Entropy Masking for Safe Chain-of-Thought Distillation

Ziyang Ma, Qingyue Yuan, Linhai Zhang, Deyu Zhou

专题命中 安全训练 :alignment(abstract);safety(abstract);分类 cs.CL

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11504 2025-08-18 cs.LG cs.CY 62%

Predicting and Explaining Traffic Crash Severity Through Crash Feature Selection

Andrea Castellani, Zacharias Papadovasilakis, Giorgos Papoutsoglou, Mary Cole, Brian Bautsch, Tobias Rodemann, Ioannis Tsamardinos, Angela Harden

机构 * Honda Research Institute Europe(霍恩达欧洲研究机构) The Ohio State University(俄亥俄州立大学) Department of Computer Science, University of Crete(克里特大学计算机科学系) American Honda Motor Co., Inc.(美国本田摩托公司)

专题命中 安全训练 :safety(abstract);分类 cs.CY、cs.LG

Comments Preprint. Manuscript under review at "Accident Analysis & Prevention" journal

详情

展开后加载摘要…

URL PDF HTML 收藏