arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-08-08 至 2025-08-08 共收录 4 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全训练 4 篇

2508.05360 2025-08-08 cs.CY cs.AI 81%

Building Effective Safety Guardrails in AI Education Tools

Hannah-Beth Clark, Laura Benton, Emma Searle, Margaux Dowland, Matthew Gregory, Will Gayne, John Roberts

专题命中 安全训练 :safety(title,abstract);分类 cs.AI、cs.CY

Comments 9 pages, published in proceedings of International Conference on Artificial Intelligence in Education (AIED) 2025: Posters and Late Breaking Results, Workshops and Tutorials, Industry and Innovation Tracks, Practitioners, Doctoral Consortium, Blue Sky, and WideAIED

Journal ref Communications in Computer and Information Science (2025) vol 2590, pp. 129-136

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.04377 2025-08-08 cs.CL 79%

PolyGuard: A Multilingual Safety Moderation Tool for 17 Languages

Priyanshu Kumar, Devansh Jain, Akhila Yerukola, Liwei Jiang, Himanshu Beniwal, Thomas Hartvigsen, Maarten Sap

专题命中 安全训练 :safety(title,abstract);分类 cs.CL

Comments Accepted to COLM 2025 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05527 2025-08-08 cs.CV 78%

AI vs. Human Moderators: A Comparative Evaluation of Multimodal LLMs in Content Moderation for Brand Safety

Adi Levi, Or Levi, Sardhendu Mishra, Jonathan Morra

机构 * Zefr Inc(Zefr公司)

专题命中 安全训练 :safety(title,abstract)

Comments Accepted to the Computer Vision in Advertising and Marketing (CVAM) workshop at ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04755 2025-08-08 cs.LG cs.CE 57%

Are Large Language Models Dynamic Treatment Planners? An In Silico Study from a Prior Knowledge Injection Angle

Zhiyao Luo, Tingting Zhu

机构 * Department of Engineering Science University of Oxford(工程科学系牛津大学)

专题命中 安全训练 :safety(abstract);分类 cs.LG

Comments 20 pages

详情

展开后加载摘要…

URL PDF HTML 收藏