arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-09-10 至 2025-09-10 共收录 3 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全训练 3 篇

2503.16833 2025-09-10 cs.SD cs.AI cs.CL cs.CY eess.AS 67%

The Model Hears You: Audio Language Model Deployments Should Consider the Principle of Least Privilege

Luxi He, Xiangyu Qi, Michel Liao, Inyoung Cheong, Prateek Mittal, Danqi Chen, Peter Henderson

专题命中 安全训练 :safety(abstract);分类 cs.CL、cs.AI、cs.CY

Comments Published at AIES 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07845 2025-09-10 cs.LG 57%

Predicting person-level injury severity using crash narratives: A balanced approach with roadway classification and natural language process techniques

Mohammad Zana Majidi, Sajjad Karimi, Teng Wang, Robert Kluger, Reginald Souleyrette

专题命中 安全训练 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07304 2025-09-10 eess.SY cs.SY 50%

Distributed Leader-Follower Consensus for Uncertain Multiagent Systems with Time-Triggered Switching of the Communication Network

Armel Koulong, Ali Pakniyat

专题命中 安全训练 :safety(abstract)

Comments Joint submission paper MECC-JDSMC. Accepted for the 2025 Modeling, Estimation and Control Conference (MECC). Currently under review by the ASME Journal of Dynamic Systems, Measurement, and Control (JDSMC)

详情

展开后加载摘要…

URL PDF HTML 收藏