arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-11-18 至 2025-11-18 共收录 7 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全训练 7 篇

2511.12982 2025-11-18 cs.CR cs.CV 88%

SafeGRPO: Self-Rewarded Multimodal Safety Alignment via Rule-Governed Policy Optimization

Xuankun Rong, Wenke Huang, Tingfeng Wang, Daiguo Zhou, Bo Du, Mang Ye

机构 * School of Computer Science, Wuhan University(武汉大学计算机学院) MiLM Plus, Xiaomi Inc.(小米公司)

专题命中 安全训练 :alignment(title,abstract);safety(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12497 2025-11-18 cs.CL cs.AI cs.CR 84%

SGuard-v1: Safety Guardrail for Large Language Models

JoonHo Lee, HyeonMin Cho, Jaewoong Yun, Hyunjae Lee, JunKyu Lee, Juree Seok

专题命中 安全训练 :safety(title,abstract);AI safety(abstract);分类 cs.CL、cs.AI

Comments Technical Report

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11693 2025-11-18 cs.AI cs.CR cs.CV cs.LG 73%

Value-Aligned Prompt Moderation via Zero-Shot Agentic Rewriting for Safe Image Generation

Xin Zhao, Xiaojun Chen, Bingshan Liu, Zeyao Liu, Zhendong Zhao, Xiaoyan Gu

专题命中 安全训练 :alignment(abstract);safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14031 2025-11-18 cs.CL cs.AI cs.LG 71%

Unintended Misalignment from Agentic Fine-Tuning: Risks and Mitigation

Dongyoon Hahm, Taywon Min, Woogyeol Jin, Kimin Lee

专题命中 安全训练 :safety(abstract,comments);分类 cs.CL、cs.AI、cs.LG;alignment(comments)

Comments Accepted at AAAI 2026 AI Alignment Track, Source code: https://github.com/HahmDY/agentic-ft-safety

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12991 2025-11-18 cs.CL 57%

Fine-Tuned LLMs Know They Don't Know: A Parameter-Efficient Approach to Recovering Honesty

Zeyu Shi, Ziming Wang, Tianyu Chen, Shiqi Gao, Haoyi Zhou, Qingyun Sun, Jianxin Li

专题命中 安全训练 :trustworthy(abstract);分类 cs.CL

Comments Accepted by AAAI 2026 Main Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09073 2025-11-18 cs.FL cs.AI cs.GT 57%

Good-for-MDP State Reduction for Stochastic LTL Planning

Christoph Weinhuber, Giuseppe De Giacomo, Yong Li, Sven Schewe, Qiyi Tang

专题命中 安全训练 :safety(abstract);分类 cs.AI

Comments 16 pages including appendices, accepted to AAAI 2026; fixed some typoes

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12160 2025-11-18 cs.RO 50%

Game-Theoretic Safe Multi-Agent Motion Planning with Reachability Analysis for Dynamic and Uncertain Environments (Extended Version)

Wenbin Mai, Minghui Liwang, Xinlei Yi, Xiaoyu Xia, Seyyedali Hosseinalipour, Xianbin Wang

机构 * Department of Electrical and Computer Engineering, National University of Singapore(国立新加坡大学电气与计算机工程系) Department of Control Science and Engineering, Shanghai Institute of Intelligent Science and Technology(上海智能科学与技术研究院控制科学与工程系)

专题命中 安全训练 :safety(abstract)

Comments 12 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏