arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 3302 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全训练 3302 篇

2512.05686 2025-12-08 eess.SY cs.SY 78%

LA-RL: Language Action-guided Reinforcement Learning with Safety Guarantees for Autonomous Highway Driving

LA-RL: 基于语言动作引导的安全强化学习用于自动驾驶高速公路驾驶

Yiming Shu, Jiahui Xu, Jiwei Tang, Ruiyang Gao, Chen Sun

专题命中 安全训练 :safety(title,abstract)

AI总结 LA-RL通过整合大语言模型的语义推理和改进的安全层,提升自动驾驶高速公路驾驶的效率与安全性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.12610 2025-11-25 eess.SY cs.SY math.OC 78%

Distributed Safe Control Design and Probabilistic Safety Verification for Multi-Agent Systems

分布式安全控制设计与多智能体系统的概率安全验证

Han Wang, Antonis Papachristodoulou, Kostas Margellos

专题命中 安全训练 :safety(title,abstract)

AI总结 本文提出了一种分布式算法,用于多智能体系统的安全控制设计和概率安全验证,通过合作机制解决不可行性问题,并通过场景方法量化安全性。

Comments manuscript accepted by Automatica

Journal ref Automatica, Vol. 179, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14433 2025-11-19 cs.LO cs.RO cs.SE 78%

Safe-ROS: An Architecture for Autonomous Robots in Safety-Critical Domains

Diana C. Benjumea, Marie Farrell, Louise A. Dennis

机构 * Department of Computer Science The University of Manchester Manchester, UK(计算机科学系曼彻斯特大学曼彻斯特英国) University of Manchester Manchester, UK(曼彻斯特大学曼彻斯特英国)

专题命中 安全训练 :safety(title,abstract)

Comments In Proceedings FMAS 2025, arXiv:2511.13245

Journal ref EPTCS 436, 2025, pp. 48-68

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.12520 2025-10-28 cs.CV 78%

SafeEraser: Enhancing Safety in Multimodal Large Language Models through Multimodal Machine Unlearning

Junkai Chen, Zhijie Deng, Kening Zheng, Yibo Yan, Shuliang Liu, PeiJun Wu, Peijie Jiang, Jia Liu, Xuming Hu

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) The Hong Kong University of Science and Technology(香港科技大学) Southeast University(东南大学) Ant Group, Alibaba(蚂蚁集团)

专题命中 安全训练 :safety(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.02293 2025-10-09 cs.RO cs.MA cs.SY eess.SY 78%

Resolving Conflicting Constraints in Multi-Agent Reinforcement Learning with Layered Safety

Jason J. Choi, Jasmine Jerry Aloor, Jingqi Li, Maria G. Mendoza, Hamsa Balakrishnan, Claire J. Tomlin

机构 * University of California, Berkeley(加州大学伯克利分校) Massachusetts Institute of Technology(麻省理工学院)

专题命中 安全训练 :safety(title,abstract)

Comments Accepted for publication at the 2025 Robotics: Science and Systems Conference. 18 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02363 2025-10-06 eess.SY cs.SY 78%

Precise HDV Positioning through Safety-Aware Integrated Sensing and Communication in a Value-of-Information-Driven 6G V2X System

Mohammad Reza Abedi, Zahra Rashidi, Nader Mokari, Hamid Saeedi, Nizar Zorba

专题命中 安全训练 :safety(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24418 2025-09-30 cs.CR 78%

GSPR: Aligning LLM Safeguards as Generalizable Safety Policy Reasoners

Haoran Li, Yulin Chen, Jingru Zeng, Hao Peng, Huihao Jing, Wenbin Hu, Xi Yang, Ziqian Zeng, Sirui Han, Yangqiu Song

专题命中 安全训练 :safety(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.13164 2025-09-18 cs.RO cs.SY eess.SY 78%

TeraSim-World: Worldwide Safety-Critical Data Synthesis for End-to-End Autonomous Driving

Jiawei Wang, Haowei Sun, Xintao Yan, Shuo Feng, Jun Gao, Henry X. Liu

机构 * University of Michigan(密歇根大学) SaferDrive AI The University of Hong Kong(香港大学) Tsinghua University(清华大学) NVIDIA(英伟达)

专题命中 安全训练 :safety(title,abstract)

Comments 8 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.03486 2025-09-12 cs.CR cs.CV cs.SI 78%

UnsafeBench: Benchmarking Image Safety Classifiers on Real-World and AI-Generated Images

Yiting Qu, Xinyue Shen, Yixin Wu, Michael Backes, Savvas Zannettou, Yang Zhang

机构 * CISPA Helmholtz Center for Information Security(CISPA赫尔姆霍茨信息安全中心) TU Delft(代尔夫特理工大学)

专题命中 安全训练 :safety(title,abstract)

Comments To Appear in the ACM Conference on Computer and Communications Security (CCS), October 13, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.09423 2025-09-09 cs.RO cs.CV 78%

Efficient Alignment of Unconditioned Action Prior for Language-conditioned Pick and Place in Clutter

Kechun Xu, Xunlong Xia, Kaixuan Wang, Yifei Yang, Yunxuan Mao, Bing Deng, Jieping Ye, Rong Xiong, Yue Wang

机构 * Zhejiang University and Alibaba Cloud(浙江大学和阿里云) Alibaba Cloud(阿里云) Zhejiang University(浙江大学)

专题命中 安全训练 :alignment(title,abstract)

Comments Accepted by T-ASE and CoRL25 GenPriors Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05527 2025-08-08 cs.CV 78%

AI vs. Human Moderators: A Comparative Evaluation of Multimodal LLMs in Content Moderation for Brand Safety

Adi Levi, Or Levi, Sardhendu Mishra, Jonathan Morra

机构 * Zefr Inc(Zefr公司)

专题命中 安全训练 :safety(title,abstract)

Comments Accepted to the Computer Vision in Advertising and Marketing (CVAM) workshop at ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.09845 2025-08-05 cs.RO 78%

Friction-Aware Safety Locomotion for Wheeled-legged Robots using Vision Language Models and Reinforcement Learning

Bo Peng, Donghoon Baek, Qijie Wang, Joao Ramos

机构 * University of Illinois at Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Department of Electrical and Computer Engineering(电气与计算机工程系) Department of Mechanical Science Engineering(机械科学与工程系) School of Software(软件学院)

专题命中 安全训练 :safety(title,abstract)

Comments Accepted to Humanoids 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.23226 2025-08-01 cs.CV 78%

Toward Safe, Trustworthy and Realistic Augmented Reality User Experience

Yanming Xiu

机构 * Department of Electrical and Computer Engineering, Duke University(电子工程系,杜克大学)

专题命中 安全训练 :trustworthy(title);safety(abstract)

Comments 2 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20685 2025-07-29 eess.SY cs.SY 78%

What's Really Different with AI? -- A Behavior-based Perspective on System Safety for Automated Driving Systems

Marcus Nolte, Nayel Fabian Salem, Olaf Franke, Jan Heckmann, Christoph Höhmann, Georg Stettinger, Markus Maurer

专题命中 安全训练 :safety(title,abstract)

Comments 8 pages, 1 figure, 1 table, to be published in 2025 IEEE International Automated Vehicle Validation Conference (IAVVC)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.04980 2025-07-16 cs.RO cs.SY eess.SY 78%

LVLM-MPC Collaboration for Autonomous Driving: A Safety-Aware and Task-Scalable Control Architecture

Kazuki Atsuta, Kohei Honda, Hiroyuki Okuda, Tatsuya Suzuki

机构 * Department of Mechanical Systems Engineering(机械系统工程系) Nagoya University(名古屋大学)

专题命中 安全训练 :safety(title,abstract)

Comments 8 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.22894 2025-07-01 cs.RO 78%

Safe Reinforcement Learning with a Predictive Safety Filter for Motion Planning and Control: A Drifting Vehicle Example

Bei Zhou, Baha Zarrouki, Mattia Piccinini, Cheng Hu, Lei Xie, Johannes Betz

机构 * State Key Laboratory of Industrial Control Technology, Zhejiang University(浙江大学工业控制技术国家重点实验室) Professorship of Autonomous Vehicle Systems, Technical University of Munich(慕尼黑工业大学自主车辆系统教授职位)

专题命中 安全训练 :safety(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.02620 2025-06-04 eess.SY cs.RO cs.SY 78%

Back to Base: Towards Hands-Off Learning via Safe Resets with Reach-Avoid Safety Filters

Azra Begzadić, Nikhil Uday Shinde, Sander Tonkens, Dylan Hirsch, Kaleb Ugalde, Michael C. Yip, Jorge Cortés, Sylvia Herbert

专题命中 安全训练 :safety(title,abstract)

Comments The first three authors contributed equally to the work

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19860 2025-05-27 cs.RO 78%

Causal Bayesian Networks for Data-driven Safety Analysis of Complex Systems

Roman Gansch, Lina Putze, Tjark Koopmann, Jan Reich, Christian Neurohr

机构 * Robert Bosch GmbH, Corporate Research(罗伯特·博世有限公司,企业研究) German Aerospace Center (DLR) e.V., Institute of Systems Engineering for Future Mobility(德国航空航天中心(DLR)协会,未来交通系统工程研究所) Fraunhofer Institute for Experimental Software Engineering (IESE)(弗劳恩霍夫实验软件工程研究所)

专题命中 安全训练 :safety(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.00973 2025-05-06 eess.SY cs.SY 78%

Defense Strategies for Autonomous Multi-agent Systems: Ensuring Safety and Resilience Under Exponentially Unbounded FDI Attacks

Yichao Wang, Mohamadamin Rajabinezhad, Dimitra Panagou, Shan Zuo

专题命中 安全训练 :safety(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.04153 2025-04-21 cs.RO math.OC 78%

A Dynamic Safety Shield for Safe and Efficient Reinforcement Learning of Navigation Tasks

Murad Dawood, Ahmed Shokry, Maren Bennewitz

专题命中 安全训练 :safety(title,abstract)

Comments Accepted in L4DC2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.06670 2025-04-10 cs.RO 78%

Dynamic Residual Safe Reinforcement Learning for Multi-Agent Safety-Critical Scenarios Decision-Making

Kaifeng Wang, Yinsong Chen, Qi Liu, Xueyuan Li, Xin Gao

专题命中 安全训练 :safety(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.14656 2025-03-20 cs.RO math.OC 78%

Safety-Critical and Distributed Nonlinear Predictive Controllers for Teams of Quadrupedal Robots

Basit Muhammad Imran, Jeeseop Kim, Taizoon Chunawala, Alexander Leonessa, Kaveh Akbari Hamed

专题命中 安全训练 :safety(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.06892 2025-03-11 cs.RO 78%

SafePlan: Leveraging Formal Logic and Chain-of-Thought Reasoning for Enhanced Safety in LLM-based Robotic Task Planning

Ike Obi, Vishnunandan L. N. Venkatesh, Weizheng Wang, Ruiqi Wang, Dayoon Suh, Temitope I. Amosa, Wonse Jo, Byung-Cheol Min

专题命中 安全训练 :safety(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.13697 2025-01-24 eess.SY cs.SY stat.ML 78%

Safety in safe Bayesian optimization and its ramifications for control

Christian Fiedler, Johanna Menn, Sebastian Trimpe

专题命中 安全训练 :safety(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.16809 2025-01-08 cs.CE 78%

Analytical assessment of workers' safety concerning direct and indirect ways of getting infected by dangerous pathogen

Krzysztof Domino, Arkadiusz Sochan, Jarosław Adam Miszczak

专题命中 安全训练 :safety(title,abstract)

Journal ref Journal of Computational Science Volume 85, February 2025, 102509

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.15437 2024-12-23 eess.SY cs.SY 78%

Safety-Critical Control of Discontinuous Systems with Nonsmooth Safe Sets

Mohammed Alyaseen, Nikolay Atanasov, Jorge Cortes

专题命中 安全训练 :safety(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.10031 2024-11-18 eess.SY cs.SY 78%

Enforcing Cooperative Safety for Reinforcement Learning-based Mixed-Autonomy Platoon Control

Jingyuan Zhou, Longhao Yan, Jinhao Liang, Kaidi Yang

专题命中 安全训练 :safety(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.03570 2024-11-04 cs.RO 78%

Embodied AI with Two Arms: Zero-shot Learning, Safety and Modularity

Jake Varley, Sumeet Singh, Deepali Jain, Krzysztof Choromanski, Andy Zeng, Somnath Basu Roy Chowdhury, Avinava Dubey, Vikas Sindhwani

专题命中 安全训练 :safety(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.03885 2024-10-08 cs.RO cs.SY eess.SY math.OC 78%

Collaborative Safety-Critical Formation Control with Obstacle Avoidance

Brooks A. Butler, Chi Ho Leung, Philip E. Paré

专题命中 安全训练 :safety(title,abstract)

Comments This work is under review for publication in Automatica. arXiv admin note: text overlap with arXiv:2311.11156

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.15343 2024-09-25 cs.IR 78%

Advertiser Content Understanding via LLMs for Google Ads Safety

Joseph Wallace, Tushar Dogra, Wei Qiao, Yuan Wang

专题命中 安全训练 :safety(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏