arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 3317 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全训练 3317 篇

2508.11425 2025-11-25 cs.MA 50%

Tapas Are Free! Training-Free Adaptation of Programmatic Agents via LLM-Guided Program Synthesis in Dynamic Environments

Tapas Are Free! 通过LLM引导的程序合成实现动态环境下的免训练程序代理适应

Jinwei Hu, Yi Dong, Youcheng Sun, Xiaowei Huang

专题命中 安全训练 :safety(abstract)

AI总结 TAPA通过LLM引导的程序合成,在动态环境中实现免训练的程序代理适应,提升网络安全和群体智能任务的性能与可靠性。

Comments Extended version of the paper accepted at AAAI-26 Oral to comply with the AAAI camera-ready requirements with minor revisions

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.16958 2025-11-24 econ.TH 50%

Real Option AI: Reversibility, Silence, and the Release Ladder

实时选项AI:可逆性、沉默与释放阶梯

I. Sebastian Buhai

专题命中 安全训练 :safety(abstract)

AI总结 本文研究了AI产品发布节奏作为战略实时选项的最优执行,揭示了可逆性、沉默与释放阶梯的内生机制,以及杠杆对不可逆性的影响。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15358 2025-11-20 cs.RO 50%

Platform-Agnostic Reinforcement Learning Framework for Safe Exploration of Cluttered Environments with Graph Attention

Gabriele Calzolari, Vidya Sumathy, Christoforos Kanellakis, George Nikolakopoulos

机构 * Robotics and AI Group, Department of Computer Science, Electrical and Space Engineering, Luleå University of Technology(机器人与人工智能组,计算机科学、电气与空间工程系,吕勒奥技术大学)

专题命中 安全训练 :safety(abstract)

Comments 8 pages, 6 figures, submitted to the 2026 IEEE International Conference on Robotics & Automation

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14378 2025-11-19 physics.soc-ph 50%

Emergent Cooperative Driving Strategies for Stop-and-Go Wave Mitigation via Multi-Agent Reinforcement Learning

Raphael Korbmacher, Daniel Straub, Antoine Tordeux, Claudia Totzeck

专题命中 安全训练 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11567 2025-11-19 eess.SY cs.RO cs.SY 50%

Who Moved My Distribution? Conformal Prediction for Interactive Multi-Agent Systems

Allen Emmanuel Binny, Anushri Dixit

专题命中 安全训练 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.12274 2025-11-19 cs.RO 50%

Hierarchical LLMs In-the-Loop Optimization for Real-Time Multi-Robot Target Tracking under Unknown Hazards

Yuwei Wu, Yuezhan Tao, Peihan Li, Guangyao Shi, Gaurav S. Sukhatme, Vijay Kumar, Lifeng Zhou

机构 * GRASP Lab, University of Pennsylvania(宾夕法尼亚大学GRASP实验室) Department of Electrical and Computer Engineering, Drexel University(德雷塞尔大学电气与计算机工程系) Department of Computer Science, University of Southern California(南加州大学计算机科学系)

专题命中 安全训练 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12160 2025-11-18 cs.RO 50%

Game-Theoretic Safe Multi-Agent Motion Planning with Reachability Analysis for Dynamic and Uncertain Environments (Extended Version)

Wenbin Mai, Minghui Liwang, Xinlei Yi, Xiaoyu Xia, Seyyedali Hosseinalipour, Xianbin Wang

机构 * Department of Electrical and Computer Engineering, National University of Singapore(国立新加坡大学电气与计算机工程系) Department of Control Science and Engineering, Shanghai Institute of Intelligent Science and Technology(上海智能科学与技术研究院控制科学与工程系)

专题命中 安全训练 :safety(abstract)

Comments 12 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10586 2025-11-14 eess.SY cs.RO cs.SY 50%

Safe Planning in Interactive Environments via Iterative Policy Updates and Adversarially Robust Conformal Prediction

Omid Mirzaeedodangeh, Eliot Shekhtman, Nikolai Matni, Lars Lindemann

机构 * Automatic Control Laboratory (IfA)(自动控制实验室)

专题命中 安全训练 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09813 2025-11-14 cs.HC 50%

I've Seen Enough: Measuring the Toll of Content Moderation on Mental Health

Gabrielle M Gauthier, Eesha Ali, Amna Asim, Sarah Cornell-Maier, Lori A. Zoellner

专题命中 安全训练 :red teaming(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09013 2025-11-13 cs.RO cs.CV 50%

UniMM-V2X: MoE-Enhanced Multi-Level Fusion for End-to-End Cooperative Autonomous Driving

Ziyi Song, Chen Xia, Chenbing Wang, Haibao Yu, Sheng Zhou, Zhisheng Niu

专题命中 安全训练 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16240 2025-11-04 cs.RO 50%

Cosmos-Surg-dVRK: World Foundation Model-based Automated Online Evaluation of Surgical Robot Policy Learning

Lukas Zbinden, Nigel Nelson, Juo-Tung Chen, Xinhao Chen, Ji Woong Kim, Mahdi Azizian, Axel Krieger, Sean Huver

机构 * NVIDIA Johns Hopkins University(约翰霍普金斯大学) Stanford University(斯坦福大学)

专题命中 安全训练 :alignment(abstract)

Comments minor metadata and notation fixes; +3 citations

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23899 2025-10-29 cs.MA cs.RO 50%

Coordinated Autonomous Drones for Human-Centered Fire Evacuation in Partially Observable Urban Environments

Maria G. Mendoza, Addison Kalanther, Daniel Bostwick, Emma Stephan, Chinmay Maheshwari, Shankar Sastry

机构 * Mechanical Engineering University of California, Berkeley(机械工程 加州大学伯克利分校) Computer Sciences University of California, Berkeley(计算机科学 加州大学伯克利分校) Computer Engineering Johns Hopkins University(计算机工程 约翰霍普金斯大学)

专题命中 安全训练 :safety(abstract)

Comments Accepted to IEEE Global Humanitarian Technology Conference (GHTC 2025). 8 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.13939 2025-10-28 cs.CV 50%

Med-R1: Reinforcement Learning for Generalizable Medical Reasoning in Vision-Language Models

Yuxiang Lai, Jike Zhong, Ming Li, Shitian Zhao, Yuheng Li, Konstantinos Psounis, Xiaofeng Yang

机构 * Department of Computer Science and Informatics, Emory University(计算机科学与信息学系,埃默里大学) Department of Computer Science and Department of Electrical and Computer Engineering, University of Southern California(计算机科学系和电气与计算机工程系,南加州大学) Department of Computer Science, University of Tokyo(计算机科学系,东京大学) Department of Computer Science, Johns Hopkins University(计算机科学系,约翰霍普金斯大学) Department of Biomedical Engineering, Georgia Institute of Technology and Emory University(生物医学工程系,佐治亚理工学院和埃默里大学)

专题命中 安全训练 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15937 2025-10-21 q-fin.RM q-fin.TR 50%

Tail-Safe Stochastic-Control SPX-VIX Hedging: A White-Box Bridge Between AI Sensitivities and Arbitrage-Free Market Dynamics

Jian'an Zhang

专题命中 安全训练 :safety(abstract)

Comments 52 pages; 3 figures; PRIMEarxiv template; fully reproducible artifact (code, configs, plots)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.00682 2025-10-21 cs.RO 50%

Immersive Explainability: Visualizing Robot Navigation Decisions through XAI Semantic Scene Projections in Virtual Reality

Jorge de Heuvel, Sebastian Müller, Marlene Wessels, Aftab Akhtar, Christian Bauckhage, Maren Bennewitz

机构 * University of Bonn(波恩大学) Lamarr Institute for Machine Learning and Artificial Intelligence(拉马尔人工智能与机器学习研究所) Center for Robotics(机器人中心) University of Mainz(美因茨大学) Fraunhofer Institute for Intelligent Analysis and Information Systems IAIS(弗劳恩霍夫智能分析与信息系统研究所)

专题命中 安全训练 :trustworthy(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.12861 2025-10-21 cs.RO 50%

Safe Multi-Agent Reinforcement Learning for Behavior-Based Cooperative Navigation

Murad Dawood, Sicong Pan, Nils Dengler, Siqi Zhou, Angela P. Schoellig, Maren Bennewitz

机构 * Humanoid Robots Lab, University of Bonn(波恩大学人形机器人实验室)

专题命中 安全训练 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.02679 2025-10-20 eess.SY cs.SY 50%

A Set-Theoretic Robust Control Approach for Linear Quadratic Games with Unknown Counterparts

Francesco Bianchin, Robert Lefringhausen, Elisa Gaetan, Samuel Tesfazgi, Sandra Hirche

专题命中 安全训练 :safety(abstract)

Comments Accepted for publication in the Proceedings of the 64th IEEE Conference on Decision and Control

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12477 2025-10-15 cs.RO 50%

A Task-Efficient Reinforcement Learning Task-Motion Planner for Safe Human-Robot Cooperation

Gaoyuan Liu, Joris de Winter, Kelly Merckaert, Denis Steckelmacher, Ann Nowe, Bram Vanderborght

机构 * Department of Mechanical Engineering, Vrije Universiteit Brussel(布鲁塞尔自由大学机械工程系) imec Flanders Make(弗拉芒制造) Artificial Intelligence (AI) Lab, Vrije Universiteit Brussel(布鲁塞尔自由大学人工智能实验室)

专题命中 安全训练 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11185 2025-10-14 cs.HC 50%

Principles of Safe AI Companions for Youth: Parent and Expert Perspectives

Yaman Yu, Mohi, Aishi Debroy, Xin Cao, Karen Rudolph, Yang Wang

专题命中 安全训练 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08917 2025-10-13 cs.HC 50%

"I know it's not right, but that's what it said to do": Investigating Trust in AI Chatbots for Cybersecurity Policy

Brandon Lit, Edward Crowder, Daniel Vogel, Hassan Khan

专题命中 安全训练 :prompt injection(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.07604 2025-10-10 cs.SE 50%

RustAssure: Differential Symbolic Testing for LLM-Transpiled C-to-Rust Code

Yubo Bai, Tapti Palit

专题命中 安全训练 :safety(abstract)

Comments 13 pages to appear in Proceedings of ASE 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.07206 2025-10-09 cs.CV 50%

EigenScore: OOD Detection using Covariance in Diffusion Models

Shirin Shoushtari, Yi Wang, Xiao Shi, M. Salman Asif, Ulugbek S. Kamilov

机构 * Washington University in St. Louis(华盛顿大学圣路易斯分校) University of California, Riverside(加州大学河滨分校)

专题命中 安全训练 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06666 2025-10-09 math.OC 50%

Trajectory-Optimized Density Control with Flow Matching

Xu Duan, Dongmei Chen

专题命中 安全训练 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21292 2025-10-07 cs.SE 50%

Semantic Clustering of Civic Proposals: A Case Study on Brazil's National Participation Platform

Ronivaldo Ferreira, Guilherme da Silva, Carla Rocha, Gustavo Pinto

专题命中 安全训练 :alignment(abstract)

Comments 12 pages, in Portuguese language

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04076 2025-10-07 cs.RO cs.SY eess.SY 50%

From Shadow to Light: Toward Safe and Efficient Policy Learning Across MPC, DeePC, RL, and LLM Agents

Amin Vahidi-Moghaddam, Sayed Pedram Haeri Boroujeni, Iman Jebellat, Ehsan Jebellat, Niloufar Mehrabi, Zhaojian Li

专题命中 安全训练 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.00882 2025-10-07 cs.SE 50%

SAFE: Advancing Large Language Models in Leveraging Semantic and Syntactic Relationships for Software Vulnerability Detection

Van Nguyen, Surya Nepal, Tingmin Wu, Xingliang Yuan, Carsten Rudolph

专题命中 安全训练 :safety(abstract)

Journal ref Proceedings of the 20th ACM Asia Conference on Computer and Communications Security (ASIA CCS), 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01623 2025-10-03 cs.CV cs.RO 50%

VLA-R1: Enhancing Reasoning in Vision-Language-Action Models

Angen Ye, Zeyu Zhang, Boyuan Wang, Xiaofeng Wang, Dapeng Zhang, Zheng Zhu

机构 * GigaAI CASIA Tsinghua University(清华大学)

专题命中 安全训练 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.16144 2025-10-02 eess.SY cs.SY 50%

Safe Event-triggered Gaussian Process Learning for Barrier-Constrained Control

Armin Lederer, Azra Begzadić, Sandra Hirche, Jorge Cortés, Sylvia Herbert

专题命中 安全训练 :safety(abstract)

Comments The first two authors contributed equally to the work

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.26593 2025-10-01 cs.HC 50%

Exploring Large Language Model as an Interactive Sports Coach: Lessons from a Single-Subject Half Marathon Preparation

Kichang Lee

专题命中 安全训练 :safety(abstract)

Comments 23 pages, 21 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21102 2025-09-26 cs.CV 50%

Mammo-CLIP Dissect: A Framework for Analysing Mammography Concepts in Vision-Language Models

Suaiba Amina Salahuddin, Teresa Dorszewski, Marit Almenning Martiniussen, Tone Hovda, Antonio Portaluri, Solveig Thrun, Michael Kampffmeyer, Elisabeth Wetzer, Kristoffer Wickstrøm, Robert Jenssen

机构 * UiT The Arctic University of Norway(乌塔大学极地大学) Technical University of Denmark(技术大学) Østfold Hospital Trust(奥斯fold医院信托) Vestre Viken Hospital Trust(维斯特维肯医院信托) Radboud University Nijmegen Medical Centre(拉德堡德大学奈梅亨医疗中心) The Netherlands Cancer Institute(荷兰癌症研究所) Antoni van Leeuwenhoek Hospital(安东尼·弗莱明医院) University of Copenhagen(哥本哈根大学)

专题命中 安全训练 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏