arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-09-08 至 2025-09-08 共收录 5 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全训练 5 篇

2509.04512 2025-09-08 cs.CL cs.LG 84%

Scaling behavior of large language models in emotional safety classification across sizes and tasks

Edoardo Pinzuti, Oliver Tüscher, André Ferreira Castro

机构 * Leibniz Institute for Resilience Research(莱比锡韧性研究所) University Medical Center Halle(哈雷医学院) German Center for Mental Health (DZPG)(德国心理健康中心(DZPG)) University Medical Center of the Johannes Gutenberg-University Mainz(美因茨约瑟夫·冯·拉贝大学医学院)

专题命中 安全训练 :safety(title,abstract);alignment(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.12391 2025-09-08 cs.LG 79%

Don't Trade Off Safety: Diffusion Regularization for Constrained Offline RL

Junyu Guo, Zhi Zheng, Donghao Ying, Ming Jin, Shangding Gu, Costas Spanos, Javad Lavaei

机构 * University of California Berkeley(加州大学伯克利分校) Virginia Tech(弗吉尼亚理工大学)

专题命中 安全训练 :safety(title,abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04535 2025-09-08 cs.RO cs.AI cs.LG 62%

In-Context Policy Adaptation via Cross-Domain Skill Diffusion

Minjong Yoo, Woo Kyung Kim, Honguk Woo

专题命中 安全训练 :alignment(abstract);分类 cs.AI、cs.LG

Comments 9 pages

Journal ref Proceedings of the AAAI Conference on Artificial Intelligence, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04478 2025-09-08 cs.CL 57%

An End-to-End System for Culturally-Attuned Driving Feedback using a Dual-Component NLG Engine

Iniakpokeikiye Peter Thompson, Yi Dewei, Reiter Ehud

机构 * Dept. of Computing Science University of Aberdeen(计算科学系阿伯丁大学)

专题命中 安全训练 :safety(abstract);分类 cs.CL

Comments The paper has 5 figures and 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04714 2025-09-08 cs.SI 50%

ThumbnailTruth: A Multi-Modal LLM Approach for Detecting Misleading YouTube Thumbnails Across Diverse Cultural Settings

Wajiha Naveed, Zartash Afzal Uzmi, Zafar Ayyub Qazi

专题命中 安全训练 :trustworthy(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏