arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-11-17 至 2025-11-17 共收录 2 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. AI治理与伦理 2 篇

2503.16851 2025-11-17 cs.CR cs.CL 70%

Interpretable LLM Guardrails via Sparse Representation Steering

Zeqing He, Zhibo Wang, Huiyu Xu, Hejun Lin, Wenhui Zhang, Zhixuan Chu

机构 * The State Key Laboratory of Blockchain and Data Security, Zhejiang University, China(区块链与数据安全国家重点实验室,浙江大学,中国) School of Cyber Science and Technology, Zhejiang University, China(网络安全与技术学院,浙江大学,中国) College of Computer and Information Sciences, Fujian Agriculture and Forestry University, China(计算机与信息科学学院,福建农林大学,中国)

专题命中 AI治理与伦理 :alignment(abstract);safety(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.22896 2025-11-17 physics.soc-ph 50%

Modelling vehicle and pedestrian collective dynamics: Challenges and advances

Antoine Tordeux, Cécile Appert-Rolland, Alexandre Nicolas, Armin Seyfried, Denis Ullmo

专题命中 AI治理与伦理 :safety(abstract)

Comments 20 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏