arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2026-02-12 至 2026-02-12 共收录 8 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全训练 8 篇

2602.10161 2026-02-12 cs.CR cs.AI cs.CL 87%

Omni-Safety under Cross-Modality Conflict: Vulnerabilities, Dynamics Mechanisms and Efficient Alignment

跨模态冲突下的全方位安全:漏洞、动态机制和高效对齐

Kun Wang, Zherui Li, Zhenhong Zhou, Yitong Zhang, Yan Mi, Kun Yang, Yiming Zhang, Junhao Dong, Zhongxiang Sun, Qiankun Li, Yang Liu

机构 * Nanyang Technological University(南洋理工大学) Beijing University of Posts and Telecommunications(北京邮电大学) Tsinghua University(清华大学) Fudan University(复旦大学) University of Science and Technology of China(中国科学技术大学) Renmin University of China(中国人民大学)

专题命中 安全训练 :safety(title,abstract);alignment(title);分类 cs.CL、cs.AI

AI总结 本文提出OmniSteer方法,通过提取黄金拒绝向量和轻量级适配器提升多模态模型的安全性与通用能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.11018 2026-02-12 cs.LG cs.AI stat.ML 81%

OSIL: Learning Offline Safe Imitation Policies with Safety Inferred from Non-preferred Trajectories

OSIL: 通过非偏好轨迹推断学习离线安全模仿策略

Returaj Burnwal, Nirav Pravinbhai Bhatt, Balaraman Ravindran

机构 * Indian Institute of Technology Madras(印度理工学院马德拉斯学院)

专题命中 安全训练 :safety(title,abstract);分类 cs.AI、cs.LG

AI总结 OSIL通过非偏好轨迹推断安全策略,从离线演示中学习安全且奖励最大化的策略,优于基线方法。

Comments 21 pages, Accepted at AAMAS 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18082 2026-02-12 cs.LG cs.RO cs.SY eess.SY 79%

Provably Optimal Reinforcement Learning under Safety Filtering

在安全过滤下可证明最优的强化学习

Donggeon David Oh, Duy P. Nguyen, Haimin Hu, Jaime F. Fisac

专题命中 安全训练 :safety(title,abstract);分类 cs.LG

AI总结 本文证明在安全过滤下强化学习可实现最优性能,通过安全关键MDP和过滤MDP的理论框架,展示安全与性能优化的分离,并验证了在安全过滤器下的训练和部署方法。

Comments Accepted for publication in the proceedings of The International Association for Safe & Ethical AI (IASEAI) 2026; 17 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.10740 2026-02-12 cs.CL 79%

Reinforced Curriculum Pre-Alignment for Domain-Adaptive VLMs

强化课程预对齐用于领域自适应视觉-语言模型

Yuming Yan, Shuo Yang, Kai Tang, Sihong Chen, Yang Zhang, Ke Xu, Dan Hu, Qun Yu, Pengfei Hu, Edith C. H. Ngai

机构 * Tencent(腾讯)

专题命中 安全训练 :alignment(title,abstract);分类 cs.CL

AI总结 本文提出RCPA方法,通过课程意识的渐进调节机制,在领域自适应中平衡领域知识获取与通用能力保持。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.10458 2026-02-12 cs.AI cs.LG 62%

Found-RL: foundation model-enhanced reinforcement learning for autonomous driving

Found-RL: 基于基础模型的强化学习用于自动驾驶

Yansong Qu, Zihao Sheng, Zilin Huang, Jiancong Chen, Yuhao Luo, Tianyi Wang, Yiheng Feng, Samuel Labi, Sikai Chen

专题命中 安全训练 :alignment(abstract);分类 cs.AI、cs.LG

AI总结 Found-RL通过异步批量推理和多样化监督机制,提升自动驾驶中强化学习的效率与实时性。

Comments 39 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.11076 2026-02-12 eess.SY cs.AI cs.SY eess.SP 57%

Interpretable Attention-Based Multi-Agent PPO for Latency Spike Resolution in 6G RAN Slicing

可解释的基于注意力的多智能体PPO用于6G RAN切片中的延迟尖峰解决

Kavan Fatehi, Mostafa Rahmani Ghourtani, Amir Sonee, Poonam Yadav, Alessandra M Russo, Hamed Ahmadi, Radu Calinescu

机构 * University of York, UK(约克大学)

专题命中 安全训练 :trustworthy(abstract);分类 cs.AI

AI总结 本文提出AE-MAPPO,通过整合六个注意力机制,实现6G RAN切片中延迟尖峰的快速诊断与高效解决,兼具SLA合规性和可解释性。

Comments This work has been accepted to appear in the IEEE International Conference on Communications (ICC)

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.10793 2026-02-12 cs.LG cs.RO 57%

Semi-Supervised Cross-Domain Imitation Learning

半监督跨领域模仿学习

Li-Min Chu, Kai-Siang Ma, Ming-Hong Chen, Ping-Chun Hsieh

机构 * Department of Computer Science, National Yang Ming Chiao Tung University(计算机科学系,国家阳明交通大学)

专题命中 安全训练 :alignment(abstract);分类 cs.LG

AI总结 本文提出半监督跨领域模仿学习方法,通过结合监督与无监督学习,实现稳定且高效的数据驱动策略学习。

Comments Published in Transactions on Machine Learning Research (TMLR)

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.10481 2026-02-12 cs.CR cs.AI cs.MA 57%

Protecting Context and Prompts: Deterministic Security for Non-Deterministic AI

保护上下文和提示:非确定性AI的确定性安全

Mohan Rajagopalan, Vinay Rao

机构 * MACAW Security, Inc.(MACAW安全公司)

专题命中 安全训练 :prompt injection(abstract);分类 cs.AI

AI总结 本研究提出认证提示和上下文技术,通过加密验证和策略代数实现LLM的预防性安全,有效防止提示注入和上下文操纵攻击。

详情

展开后加载摘要…

URL PDF HTML 收藏