arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 3317 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全训练 3317 篇

2506.21693 2025-06-30 cs.SE cs.RO 50%

The DevSafeOps Dilemma: A Systematic Literature Review on Rapidity in Safe Autonomous Driving Development and Operation

Ali Nouri, Beatriz Cabrero-Daniel, Fredrik Törner, Christian Berger

机构 * Chalmers University of Technology, Department of Computer Science(查尔姆斯理工大学计算机科学系) University of Gothenburg, Department of Computer Science(哥德堡大学计算机科学系)

专题命中 安全训练 :safety(abstract)

Comments Accepted for publication in the Journal of Systems and Software (JSS)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.16748 2025-06-23 cs.RO cs.MA 50%

A Scalable Post-Processing Pipeline for Large-Scale Free-Space Multi-Agent Path Planning with PiBT

Arjo Chakravarty, Michael X. Grey, M. A. Viraj J. Muthugala, Mohan Rajesh Elara

机构 * Intrinsic Innovation LLC ROAR Lab(ROAR 实验室) Singapore University of Technology and Design(新加坡科技设计大学)

专题命中 安全训练 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.09042 2025-06-19 cs.CV 50%

Cosmos-Drive-Dreams: Scalable Synthetic Driving Data Generation with World Foundation Models

Xuanchi Ren, Yifan Lu, Tianshi Cao, Ruiyuan Gao, Shengyu Huang, Amirmojtaba Sabour, Tianchang Shen, Tobias Pfaff, Jay Zhangjie Wu, Runjian Chen, Seung Wook Kim, Jun Gao, Laura Leal-Taixe, Mike Chen, Sanja Fidler, Huan Ling

专题命中 安全训练 :safety(abstract)

Comments Only the core contributors are listed. The full list of contributors can be found in Appendix A of this paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.14749 2025-06-18 eess.SY cs.SY 50%

Swarm-STL: A Framework for Motion Planning in Large-Scale, Multi-Swarm Systems

Shiyu Cheng, Luyao Niu, Bhaskar Ramasubramanian, Andrew Clark, Radha Poovendran

专题命中 安全训练 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.14018 2025-06-18 cs.HC 50%

"I Cannot Write This Because It Violates Our Content Policy": Understanding Content Moderation Policies and User Experiences in Generative AI Products

Lan Gao, Oscar Chen, Rachel Lee, Nick Feamster, Chenhao Tan, Marshini Chetty

专题命中 安全训练 :safety(abstract)

Comments Preprint for USENIX Security 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23549 2025-06-16 cs.SE 50%

LLM-based Property-based Test Generation for Guardrailing Cyber-Physical Systems

Khashayar Etemadi, Marjan Sirjani, Mahshid Helali Moghadam, Per Strandberg, Paul Pettersson

专题命中 安全训练 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.12952 2025-06-16 eess.SY cs.SY 50%

Safe Physics-Informed Machine Learning for Dynamics and Control

Jan Drgona, Truong X. Nghiem, Thomas Beckers, Mahyar Fazlyab, Enrique Mallada, Colin Jones, Draguna Vrabie, Steven L. Brunton, Rolf Findeisen

专题命中 安全训练 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.10279 2025-06-13 cs.RO cs.SY eess.SY 50%

Learning Safe Control via On-the-Fly Bandit Exploration

Alexandre Capone, Ryan Cosner, Aaaron Ames, Sandra Hirche

机构 * Robotics Institute, Carnegie Mellon University(卡内基梅隆大学机器人研究所) TUM School of Computation, Information and Technology, Technical University of Munich(慕尼黑技术大学计算、信息与技术学院)

专题命中 安全训练 :safety(abstract)

Comments arXiv admin note: text overlap with arXiv:2311.02133

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.07755 2025-06-10 eess.SY cs.MA cs.RO cs.SY 50%

Deep Equivariant Multi-Agent Control Barrier Functions

Nikolaos Bousias, Lars Lindemann, George Pappas

机构 * GRASP Lab, Department of Electrical & Systems Engineering, University of Pennsylvania(宾夕法尼亚大学电气与系统工程系GRASP实验室) Department of Computer Science, University of Southern California(南加州大学计算机科学系)

专题命中 安全训练 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.06077 2025-06-09 cs.RO 50%

Self driving algorithm for an active four wheel drive racecar

Gergely Bari, Laszlo Palkovics

机构 * HUMDA Lab nonprofit Kft.(HUMDA实验室非营利公司) Széchenyi István University(塞切尼伊斯特万大学)

专题命中 安全训练 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.02436 2025-06-05 cs.RO cs.SY eess.SY 50%

Safe, Out-of-Distribution-Adaptive MPC with Conformalized Neural Network Ensembles

Jose Leopoldo Contreras, Ola Shorinwa, Mac Schwager

机构 * Department of Aeronautics and Astronautics, Stanford University(航空航天工程系,斯坦福大学)

专题命中 安全训练 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.01927 2025-06-05 eess.SY cs.SY 50%

Safe Connectivity Maintenance of Underactuated Multi-Agent Networks in Dynamic Oceanic Environments

Nicolas Hoischen, Marius Wiggert, Claire J. Tomlin

专题命中 安全训练 :safety(abstract)

Comments 8 pages, Published at European Control Conference 2024 (ECC 2024) Nicolas Hoischen and Marius Wiggert contributed equally to this work

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.13140 2025-06-04 cs.RO 50%

Learn With Imagination: Safe Set Guided State-wise Constrained Policy Optimization

Yifan Sun, Feihan Li, Weiye Zhao, Rui Chen, Tianhao Wei, Changliu Liu

专题命中 安全训练 :safety(abstract)

Comments arXiv admin note: text overlap with arXiv:2306.12594

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.20796 2025-05-28 cs.HC 50%

Describe Me Something You Do Not Remember - Challenges and Risks of Exposure Design Using Generative Artificial Intelligence for Therapy of Complex Post-traumatic Disorder

Annalisa Degenhard, Stefan Tschöke, Michael Rietzler, Enrico Rukzio

专题命中 安全训练 :safety(abstract)

Comments Extended abstract for the Workshop "Generative AI and Accessibility Workshop: Surfacing Opportunities and Risks" at the CHI conference on Human Factors in Computing Systems 2025 (CHI'25) - Accepted on 4 April 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.20679 2025-05-22 cs.CE 50%

Online Prediction-Assisted Safe Reinforcement Learning for Electric Vehicle Charging Station Recommendation in Dynamically Coupled Transportation-Power Systems

Qionghua Liao, Guilong Li, Jiajie Yu, Ziyuan Gu, Wei Ma

专题命中 安全训练 :safety(abstract)

Comments Published at Transportation Research Part C: Emerging Technologies

Journal ref Transportation Research Part C 176C (2025) 105155

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.10219 2025-05-16 cs.RO 50%

Towards Safe Robot Foundation Models Using Inductive Biases

Maximilian Tölle, Theo Gruner, Daniel Palenicek, Tim Schneider, Jonas Günster, Joe Watson, Davide Tateo, Puze Liu, Jan Peters

机构 * Technical University of Darmstadt(德累斯顿技术大学) German Research Center for AI ( dfki )(人工智能研究中心) University of Oxford(牛津大学) Robotics Institute Germany ( rig )(德国机器人研究所) Centre for Cognitive Science(认知科学中心)

专题命中 安全训练 :safety(abstract)

Comments 14 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.09837 2025-05-16 cs.RO 50%

EdgeAI Drone for Autonomous Construction Site Demonstrator

Emre Girgin, Arda Taha Candan, Coşkun Anıl Zaman

机构 * Faculty of Engineering, Aerospace Engineering, Embry-Riddle Aeronautical University(工程学院,航空航天工程系,埃姆布里-里德尔航空大学) Robotics and Autonomous Systems Division, TUBITAK BILGEM(机器人与自主系统 division,TUBITAK BILGEM)

专题命中 安全训练 :safety(abstract)

Comments Paper presented at the 4th Workshop on Future of Construction: Safe, Reliable, and Precise Robots in Construction Environments, ICRA 2025, Atlanta, GA, United States

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.09144 2025-05-15 cs.RO 50%

Latent Theory of Mind: A Decentralized Diffusion Architecture for Cooperative Manipulation

Chengyang He, Gadiel Sznaier Camps, Xu Liu, Mac Schwager, Guillaume Sartoretti

机构 * National University of Singapore(新加坡国立大学) Stanford University(斯坦福大学)

专题命中 安全训练 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.07710 2025-05-13 cs.RO 50%

Hybrid Control Strategies for Safe and Adaptive Robot-Assisted Dressing

Yasmin Rafiq, Baslin A. James, Ke Xu, Robert M. Hierons, Sanja Dogramadzi

专题命中 安全训练 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.06875 2025-05-13 cs.RO 50%

Towards Human-Centric Autonomous Driving: A Fast-Slow Architecture Integrating Large Language Model Guidance with Reinforcement Learning

Chengkai Xu, Jiaqi Liu, Yicheng Guo, Yuhang Zhang, Peng Hang, Jian Sun

机构 * College of Transportation, Tongji University(同济大学交通学院)

专题命中 安全训练 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.06628 2025-05-13 cs.RO 50%

ACORN: Adaptive Contrastive Optimization for Safe and Robust Fine-Grained Robotic Manipulation

Zhongquan Zhou, Shuhao Li, Zixian Yue

机构 * School of Information Science & Engineering, Lanzhou University(信息科学与工程学院,兰州大学)

专题命中 安全训练 :safety(abstract)

Comments 6 pages,4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.04231 2025-05-08 cs.RO cs.MA cs.SY eess.SY 50%

Multi-Agent Reinforcement Learning-based Cooperative Autonomous Driving in Smart Intersections

Taoyuan Yu, Kui Wang, Zongdian Li, Tao Yu, Kei Sakaguchi

专题命中 安全训练 :safety(abstract)

Comments 7 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.11863 2025-05-08 eess.SY cs.RO cs.SY 50%

Context-aware LLM-based Safe Control Against Latent Risks

Xiyu Deng, Quan Khanh Luu, Anh Van Ho, Yorie Nakahira

专题命中 安全训练 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.00153 2025-05-02 cs.HC cs.DC 50%

Audo-Sight: Enabling Ambient Interaction For Blind And Visually Impaired Individuals

Bhanuja Ainary

专题命中 安全训练 :safety(abstract)

Comments This thesis was conducted under the guidance of Mohsen Amini Salehi. Special thanks to Minseo Kim and Jacob Bradshaw for their valuable contributions and support throughout the research process. 60 pages, 13 Figures, 2 Tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.21334 2025-05-01 cs.CV 50%

Simple Visual Artifact Detection in Sora-Generated Videos

Misora Sugiyama, Hirokatsu Kataoka

专题命中 安全训练 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.20529 2025-04-30 eess.SY cs.SY 50%

Safe Bottom-Up Flexibility Provision from Distributed Energy Resources

Costas Mylonas, Emmanouel Varvarigos, Georgios Tsaousoglou

专题命中 安全训练 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.09810 2025-04-29 cs.RO cs.SY eess.SY 50%

Think Deep and Fast: Learning Neural Nonlinear Opinion Dynamics from Inverse Dynamic Games for Split-Second Interactions

Haimin Hu, Jaime Fernández Fisac, Naomi Ehrich Leonard, Deepak Gopinath, Jonathan DeCastro, Guy Rosman

机构 * Princeton University(普林斯顿大学) Toyota Research Institute(丰田研究院)

专题命中 安全训练 :safety(abstract)

Comments IEEE International Conference on Robotics and Automation (ICRA) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.14554 2025-04-22 cs.CR cs.CV 50%

REDEditing: Relationship-Driven Precise Backdoor Poisoning on Text-to-Image Diffusion Models

Chongye Guo, Jinhu Fu, Junfeng Fang, Kun Wang, Guorui Feng

机构 * Shanghai University(上海大学) Beijing University of Posts and Telecommunications(北京邮电大学) National University of Singapore(新加坡国立大学) Nanyang Technological University(南洋理工大学)

专题命中 安全训练 :alignment(abstract)

Comments 10 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.11914 2025-04-17 cs.CV 50%

AnomalyR1: A GRPO-based End-to-end MLLM for Industrial Anomaly Detection

Yuhao Chao, Jie Liu, Jie Tang, Gangshan Wu

机构 * Nanjing University(南京大学)

专题命中 安全训练 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.15185 2025-04-10 cs.RO 50%

Semantically Safe Robot Manipulation: From Semantic Scene Understanding to Motion Safeguards

Lukas Brunke, Yanni Zhang, Ralf Römer, Jack Naimer, Nikola Staykov, Siqi Zhou, Angela P. Schoellig

专题命中 安全训练 :safety(abstract)

Comments 9 pages, 6 figures

Journal ref in IEEE Robotics and Automation Letters, vol. 10, no. 5, pp. 4810-4817, May 2025

详情

展开后加载摘要…

URL PDF HTML 收藏