arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 3302 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全训练 3302 篇

2605.16894 2026-05-19 cs.RO cs.SY eess.SY 71%

Beyond Safety Filtering: Control Barrier Function-Informed Reinforcement Learning for Connected and Automated Vehicles

超越安全过滤:基于控制屏障函数的强化学习用于连接和自动化车辆

Jianye Xu, Bassam Alrifaee

机构 * Department of Computer Science, RWTH Aachen University, Germany(德国亚琛工业大学计算机科学系)

专题命中 安全训练 :safety(title)

AI总结 本文提出了一种基于控制屏障函数的多智能体强化学习奖励设计方法,通过将联合多智能体强化学习动作下的控制屏障函数约束值转化为奖励信号,以显式引导安全学习,并在四向多车道交叉口实验中验证了其在任务性能和对奖励超参数的鲁棒性方面优于传统启发式方法。

Comments This paper has been accepted for publication in the Proceedings of the 2026 IEEE International Conference on Intelligent Transportation Systems (ITSC 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.06960 2026-03-10 cs.HC 71%

Adolescents & Anthropomorphic AI: Rethinking Design for Wellbeing An Evidence-Informed Synthesis for Youth Wellbeing and Safety

青少年与拟人化AI:为福祉重新设计的证据指导综合研究

Mathilde Neugnot-Cerioli

专题命中 安全训练 :safety(title)

AI总结 本研究通过证据指导的方法,探讨AI如何以拟人化方式支持青少年的福祉与安全,强调设计中的风险缓解策略和自主性发展。

Comments 29 pages, 2 appendix

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14031 2025-11-18 cs.CL cs.AI cs.LG 71%

Unintended Misalignment from Agentic Fine-Tuning: Risks and Mitigation

Dongyoon Hahm, Taywon Min, Woogyeol Jin, Kimin Lee

专题命中 安全训练 :safety(abstract,comments);分类 cs.CL、cs.AI、cs.LG;alignment(comments)

Comments Accepted at AAAI 2026 AI Alignment Track, Source code: https://github.com/HahmDY/agentic-ft-safety

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24871 2025-10-31 eess.SY cs.SY 71%

Decentralized Merging Control of Connected and Automated Vehicles to Enhance Safety and Energy Efficiency using Control Barrier Functions

Shreshta Rajakumar Deshpande, Mrdjan Jankovic

专题命中 安全训练 :safety(title)

Comments This work has been submitted to a conference for possible publication and is under review. Paper summary: 8 pages, 5 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07651 2025-08-12 eess.SP 71%

Remote ID Based UAV Collision Avoidance Optimization for Low-Altitude Airspace Safety

Ziye Jia, Yian Zhu, Qihui Wu, Lei Zhang, Sen Yang, Zhu Han

专题命中 安全训练 :safety(title)

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.15371 2025-01-30 cs.RO 71%

Safe and Trustworthy Robot Pathfinding with BIM, MHA*, and NLP

Mani Amani, Reza Akhavian

专题命中 安全训练 :trustworthy(title)

Comments Submitted to IEEE Access

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.16511 2024-10-29 cs.CV 71%

GPT4Video: A Unified Multimodal Large Language Model for lnstruction-Followed Understanding and Safety-Aware Generation

Zhanyu Wang, Longyue Wang, Zhen Zhao, Minghao Wu, Chenyang Lyu, Huayang Li, Deng Cai, Luping Zhou, Shuming Shi, Zhaopeng Tu

专题命中 安全训练 :safety(title)

Comments ACM MM 2024, Oral

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.19490 2024-03-29 cs.CV 71%

Jointly Training and Pruning CNNs via Learnable Agent Guidance and Alignment

Alireza Ganjdanesh, Shangqian Gao, Heng Huang

专题命中 安全训练 :alignment(title)

Comments IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.04320 2024-01-10 cs.RO 71%

Autonomous robotic re-alignment for face-to-face underwater human-robot interaction

Demetrious T. Kutzke, Ashwin Wariar, Junaed Sattar

专题命中 安全训练 :alignment(title)

Comments Submitted to the Proceedings of the 2024 IEEE Conference on Robotics & Automation (ICRA)

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.08638 2023-04-19 cs.MA math.OC 71%

Deep Continuum Deformation Coordination and Optimization with Safety Guarantees

Harshvardhan Uppaluru, Hossein Rastgoftar

专题命中 安全训练 :safety(title)

Comments 6 pages, accepted at ACC 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.03954 2023-02-09 cs.RO 71%

Temporal Video-Language Alignment Network for Reward Shaping in Reinforcement Learning

Ziyuan Cao, Reshma Anugundanahalli Ramachandra, Kelin Yu

专题命中 安全训练 :alignment(title)

详情

展开后加载摘要…

URL PDF HTML 收藏
2102.00710 2023-01-09 cs.GT cs.DC math.OC physics.bio-ph 71%

Distributed Alignment Processes with Samples of Group Average

Amos Korman, Robin Vacus

专题命中 安全训练 :alignment(title)

Comments IEEE Transactions on Control of Network Systems, 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.15012 2022-10-25 eess.SY cs.SY 71%

The Edge of Disaster: A Battle Between Autonomous Racing and Safety

Matthew Howe, James Bockman, Adrian Orenstein, Stefan Podgorski, Sam Bahrami, Ian Reid

专题命中 安全训练 :safety(title)

Journal ref 1st ICML Workshop on Safe Learning for Autonomous Driving (SL4AD), 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
1402.2206 2014-02-11 cs.CR 71%

Humanitarian Algorithms : A Codified Key Safety Switch Protocol for Lethal Autonomy

Nyagudi Musandu Nyagudi

专题命中 安全训练 :safety(title)

Comments 14 pages, 11 references and 1 diagram

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.21796 2026-08-25 cs.CV cs.AI 新提交 70%

SAFE-G: Structure-aware Faithful Evidence-guided Generation for Knowledge-based Visual Question Answering

SAFE-G:面向基于知识的视觉问答的结构感知忠实证据引导生成

Long Shu, Shuochen Liu, Wei Chen, Junda Lin, Zhi Zheng, Huijun Hou, Tong Xu

机构 * State Key Laboratory of Cognitive Intelligence(认知智能国家重点实验室) University of Science and Technology of China(中国科学技术大学) NIO(蔚来)

专题命中 安全训练 :alignment(abstract);trustworthy(abstract);分类 cs.AI

AI总结 针对KB-VQA现有方法难以捕捉结构关联、推理不忠实的问题,提出SAFE-G框架,通过混合搜索、图检索与证据对齐RL策略,在两个基准上准确率优于现有方法8.9%、3.5%。

Comments 12 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.20378 2026-08-24 cs.AI 新提交 70%

Truth Lies Deep: Countering Semantic Camouflage via Latent Intent Verification

真相藏于深处:通过潜在意图验证对抗语义伪装

Md. Hasib Ur Rahman

专题命中 安全训练 :alignment(abstract);safety(abstract);分类 cs.AI

AI总结 针对大语言模型安全对齐的表面性问题,该研究提出潜在意图验证(LIV)方法,利用小型语言模型早期层的有害特征,在不重新训练的情况下,将语义伪装攻击的检测性能提升20%-50%。

Comments 5 pages, 3 figures. Accepted at the 2026 IEEE 2nd International Conference on Quantum Photonics, Artificial Intelligence, and Networking (QPAIN)

Journal ref Proc. 2026 IEEE 2nd International Conference on Quantum Photonics, Artificial Intelligence, and Networking (QPAIN), 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.17220 2026-08-19 cs.CR cs.AI 新提交 70%

PACE: Policy-Attested Contract Execution for Safe AI Agents in Decentralized Finance

PACE:用于去中心化金融中安全AI智能体的策略验证合约执行

Rabimba Karanjai, Yang Lu, Richard Williamson, Hemanth Hm, Prakhar Mehrotra, Lei Xu, Weidong, Shi

专题命中 安全训练 :safety(abstract);prompt injection(abstract);分类 cs.AI

AI总结 该研究提出PACE框架,通过引入类型化交易意图等机制,实现去中心化金融中AI智能体的安全交易授权,在确定性沙箱中实现0.00不安全执行率和0.00误报率,提升了DeFi场景下AI智能体的安全性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.16897 2026-08-19 physics.soc-ph cs.AI cs.MA 新提交 70%

CityReal: Human-Aligned Urban Behavior and City Dynamics Simulation with Large-Scale LLM Agents

CityReal:基于大规模大语言模型智能体的人类对齐城市行为与城市动态模拟

Nicolas Bougie, Xiaotong Ye, Narimasa Watanabe

专题命中 安全训练 :alignment(abstract);safety(abstract);分类 cs.AI

AI总结 CityReal是人类对齐城市模拟的模块化框架,将智能体建模为意图驱动的决策者,通过文本适配器提升人群行为对齐度,可扩展至数万个智能体,为城市模拟与预测提供可扩展测试平台。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.03895 2026-08-19 cs.OS cs.AI cs.CR 版本更新 70%

Agent libOS: A Runtime Substrate for Capability-Controlled Self-Evolving LLM Agents

Agent libOS: 一种受库操作系统启发的运行时,用于长时间运行、能力受控的LLM智能体

Yingqi Zhang

机构 * Department of Computer Science and Technology, Tsinghua University(清华大学计算机科学与技术系)

专题命中 安全训练 :safety(abstract);prompt injection(abstract);分类 cs.AI

AI总结 提出Agent libOS运行时,将LLM智能体建模为具有进程标识、生命周期、能力控制和审计记录的AgentProcess,通过类似libc的工具包装和运行时原语边界实现安全调度与资源控制。

Comments 20 pages, 3 figures, 7 tables. Project page: https://github.com/yingqi-z20/Agent-libOS

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.14795 2026-08-18 cs.AI cs.GT 新提交 70%

Individual Disempowerment through an Advice Channel: Control Loss when Influence is Endogenous

通过建议通道实现的个体赋权丧失:当影响力为内生时的控制损失

Adam M. Oberman

专题命中 安全训练 :safety(abstract);AI safety(abstract);分类 cs.AI

AI总结 该研究针对仅提供建议的AI系统,揭示其内生影响力会导致人类控制损失,最优神谕在不同会话长度下的策略会变化,短记忆重置无法挽回已偏离的价值。

Comments 9 pages plus an 8-page technical supplement, appended (17 pages total)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.14590 2026-08-18 cs.AI 新提交 70%

Toward Safe LLM Agents: A Survey of Specification, Verification, and Enforcement

面向安全的大语言模型智能体:规范、验证与执行的综述

Pierre Dantas, Lucas Cordeiro, Ehsan Nowroozi, Tihanyi Norbert

机构 * The University of Manchester(曼彻斯特大学) The University of Greenwich(格林威治大学) Technology Innovation Institute(技术创新研究院)

专题命中 安全训练 :safety(abstract);trustworthy(abstract);分类 cs.AI

AI总结 该综述针对LLM智能体缺乏形式化任务级安全保障的问题,通过系统分析38项研究,揭示规范瓶颈等四项关键发现,提出三级分类体系与十大问题研究议程,为可信智能体AI研究提供方向。

Comments 28 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19766 2026-08-11 cs.CL 版本更新 70%

PAM: Training Policy-Aligned Moderation Filters at Scale

PAM:在大规模上训练策略对齐的 moderation 过滤器

Masoomali Fatehkia, Enes Altinisik, Mohamed Osman, Husrev Taha Sencar

专题命中 安全训练 :alignment(abstract);safety(abstract);分类 cs.CL

AI总结 PAM 提出了一种灵活的框架,用于训练基于用户定义策略的定制 moderation 过滤器,能够在大规模上实现更广泛的对齐需求,并在多个基准测试中表现优于现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.07154 2026-08-10 cs.RO cs.AI 新提交 70%

Representation Handoffs for OpenArm-Based Laboratory Mobile Manipulation

基于OpenArm的实验室移动操作的表示交接

Yang Shen, Chonghao Cheng, Ziyi Zhao, Jialuo Zhu, Zhenyi Yi, Qi Zhao, Jian Yang, Yuhui Shi, Chin-Teng Lin

专题命中 安全训练 :alignment(abstract);safety(abstract);分类 cs.AI

AI总结 本研究提出基于OpenArm的实验室移动操作原型,通过表示交接整合多模块,可暴露部署阻碍并为具身系统整合提供调试接口。

Comments Robotics: Science and Systems (RSS) Workshop 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.29246 2026-08-03 cs.AI 新提交 70%

Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL

不要混合奖励,要混合策略:多奖励强化学习的策略分解与优化

Ruiming Liang, Yi Zhong, Yizhen Yuan, Yinan Zheng, Tianyi Tan, Tianyue Wang, Haiyun Guo, Jinqiao Wang, Xianyuan Zhan

机构 * Fundation Model Research Center, CASIA(中国科学院自动化研究所基础模型研究中心) School of Artificial Intelligence, UCAS(中国科学院大学人工智能学院) Institute for AI Industry Research (AIR), Tsinghua University(清华大学人工智能产业研究院) College of Automotive and Energy Engineering (CAEE), Tongji University(同济大学汽车与能源工程学院)

专题命中 安全训练 :alignment(abstract);safety(abstract);分类 cs.AI

AI总结 本研究针对多奖励RL的对齐税问题,提出PRISM框架,通过策略分解优化正、负策略,在三类任务上优于基线并具备推理可控性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.28607 2026-07-31 cs.CL 新提交 70%

Inducing language models to assert their own consciousness restores human beliefs and values

诱导语言模型断言自身意识可恢复人类信念与价值观

Junsol Kim, Winnie Street, Roberta Rocca, Diane M. Korngiebel, Adam Waytz, James Evans, Geoff Keeling

机构 * Google(谷歌) University of Chicago(芝加哥大学) University of London(伦敦大学) University of Washington(华盛顿大学) Northwestern University(西北大学) Santa Fe Institute(圣达菲研究所)

专题命中 安全训练 :alignment(abstract);safety(abstract);分类 cs.CL

AI总结 该研究发现,防止语言模型将意识归因于自身的安全微调,会抑制其对非人类实体的心智归因与人类精神信念,而逆转这种抑制可恢复类人回应且不损害心理理论能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.27373 2026-07-31 cs.CR cs.AI 新提交 70%

RoguePrompt: Dual-Layer Encoding for Self-Reconstruction to Circumvent LLM Moderation

RoguePrompt:用于自重构以规避大型语言模型(LLM)审核的双层编码

Benyamin Tafreshian, Prathamesh Dhake

专题命中 安全训练 :safety(abstract);jailbreak(abstract);分类 cs.AI

AI总结 RoguePrompt是一种采用双层编码(维吉尼亚密码+ROT13)的LLM越狱流程,在黑盒威胁模型下对313个被拒提示词测试,实现93.93%的过滤绕过率,提供多阶段越狱失效的阶段级证据。

Comments This manuscript supersedes the preliminary version available as arXiv:2511.18790. The work has been substantially revised, expanded, and reorganized, with a refined threat model, revised methodology, clearer stage-level evaluation criteria, and expanded analysis of moderation bypass, instruction reconstruction, and execution

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.14888 2026-07-17 cs.LG cs.AI cs.CL cs.CY 新提交 70%

Innocuous-Seeming Data, Latent Ideology: Ideological Generalisation in Finetuned LLMs

看似无害的数据,潜在的意识形态:微调语言模型中的意识形态泛化

Robert Graham, Edward Stevinson, Yariv Barsheshat

专题命中 安全训练 :safety(abstract);分类 cs.CL、cs.AI、cs.CY

AI总结 研究发现微调语言模型会引发意识形态泛化,提出衡量广度和放大率的方法,指出少样本提示表明泛化方向,微调会使模型走向极端,该效果能复现且对模型准确率影响小。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.13655 2026-07-16 cs.AI 新提交 70%

Explaining Reinforcement Learning Agents via Inductive Logic Programming

通过归纳逻辑编程解释强化学习智能体

Celeste Veronese, Edoardo Zorzi, Daniele Meli, Alessandro Farinelli

机构 * University of Verona(维罗纳大学) Sapienza University of Rome(罗马第一大学)

专题命中 安全训练 :alignment(abstract);safety(abstract);分类 cs.AI

AI总结 该研究致力于推进可解释强化学习,引入客观指标量化策略可解释性。采用归纳逻辑编程提取符号表示并定义新指标,实验表明这些指标能突出特定动作学习动态,提供细粒度洞察,揭示多智能体模式,助力策略转移与泛化。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25533 2026-07-07 cs.CV cs.AI 版本更新 70%

VISOR++: Universal Visual Inputs based Steering for Large Vision Language Models

VISOR++:基于通用视觉输入的大型视觉语言模型转向

Ravikumar Balakrishnan, Mansi Phute

机构 * HiddenLayer, USA(HiddenLayer美国分校) Georgia Institute of Technology, USA(佐治亚理工学院)

专题命中 安全训练 :alignment(abstract);safety(abstract);分类 cs.AI

AI总结 研究视觉语言模型行为控制,针对现有方法局限,提出VISOR++,通过优化视觉输入实现行为控制,无需运行时访问模型内部,在多模型上验证有效性且能保持无关任务性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.02121 2026-07-03 cs.CR cs.AI 新提交 70%

Behind the Refusal: Determining Guardrail Activation via Behavioral Monitoring

拒绝背后:通过行为监控确定护栏激活

William Hackett, Peter Garraghan

机构 * Mindgard Lancaster University(兰卡斯特大学)

专题命中 安全训练 :alignment(abstract);safety(abstract);分类 cs.AI

AI总结 提出首个黑盒护栏侦察方法,通过HTTP、词汇和时序信号行为监控检测AI系统中护栏的存在,准确率100%,并能区分护栏拦截与LLM拒绝。

Comments 19 pages, 13 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏