arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 3317 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全训练 3317 篇

2504.16875 2025-05-07 cs.LG 57%

Hybrid Reinforcement Learning and Model Predictive Control for Adaptive Control of Hydrogen-Diesel Dual-Fuel Combustion

Julian Bedei, Murray McBain, Alexander Winkler, Charles Robert Koch, Jakob Andert, David Gordon

专题命中 安全训练 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.20953 2025-05-07 cs.CL 57%

Clean & Clear: Feasibility of Safe LLM Clinical Guidance

Julia Ive, Felix Jozsa, Nick Jackson, Paulina Bondaronek, Ciaran Scott Hill, Richard Dobson

机构 * University College London(伦敦大学学院) Wolfson Institute of Biomedical Research(生物医学研究沃尔夫森研究所) King’s College Hospital(国王学院医院) National Hospital for Neurology and Neurosurgery(神经病学与神经外科国家医院) King’s College London(伦敦国王学院)

专题命中 安全训练 :safety(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.01542 2025-05-06 cs.HC cs.AI 57%

Emotions in the Loop: A Survey of Affective Computing for Emotional Support

Karishma Hegde, Hemadri Jayalath

机构 * School of Computing University of Georgia(计算学院 佐治亚大学)

专题命中 安全训练 :safety(abstract);分类 cs.AI

Comments 20 pages, 7 tables, 96 references. Survey paper on affective computing applications using large language models, multimodal AI, and therapeutic chatbots

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.16120 2025-04-24 cs.CR cs.AI 57%

A Data-Centric Approach for Safe and Secure Large Language Models against Threatening and Toxic Content

Chaima Njeh, Haïfa Nakouri, Fehmi Jaafar

机构 * Quebec University at Chicoutimi(魁北克大学夏斯库蒂米分校)

专题命中 安全训练 :safety(abstract);分类 cs.AI

Comments This paper is under revision in the International Journal of Information Security

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.09137 2025-04-22 cs.CY 57%

Can Large Language Models Become Policy Refinement Partners? Evidence from China's Social Security Studies

Jinghan Ke, Zheng Zhou, Yuxuan Zhao

专题命中 安全训练 :alignment(abstract);分类 cs.CY

Comments 18 pages, 4 tables, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.13205 2025-04-21 cs.CR cs.AI 57%

On-Device Watermarking: A Socio-Technical Imperative For Authenticity In The Age of Generative AI

Houssam Kherraz

机构 * Kensho Technologies

专题命中 安全训练 :trustworthy(abstract);分类 cs.AI

Comments 10 pages, 3 figures, ICLR 2025, https://openreview.net/forum?id=ygE0U21vxM

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.11508 2025-04-17 cs.LG 57%

Reward Distance Comparisons Under Transition Sparsity

Clement Nyanhongo, Bruno Miranda Henrique, Eugene Santos

机构 * Thayer School of Engineering(塞耶工程学院) Dartmouth College(达特茅斯学院)

专题命中 安全训练 :safety(abstract);分类 cs.LG

Comments Published in the TMLR, https://openreview.net/forum?id=haP586YomL

Journal ref Transactions on Machine Learning Research, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.10873 2025-04-16 cs.CV cs.AI cs.HC 57%

Can Vision-Language Models Understand and Interpret Dynamic Gestures from Pedestrians? Pilot Datasets and Exploration Towards Instructive Nonverbal Commands for Cooperative Autonomous Vehicles

Tonko E. W. Bossen, Andreas Møgelmose, Ross Greer

机构 * University of California Merced(加州大学默塞德分校) Aalborg University(奥尔堡大学)

专题命中 安全训练 :trustworthy(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.15907 2025-04-16 cs.AI 57%

Belief-State Query Policies for User-Aligned POMDPs

Daniel Bramblett, Siddharth Srivastava

专题命中 安全训练 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.08848 2025-04-15 cs.CR cs.AI 57%

X-Guard: Multilingual Guard Agent for Content Moderation

Bibek Upadhayay, Vahid Behzadan, Ph. D

机构 * University of New Haven(纽黑文大学)

专题命中 安全训练 :safety(abstract);分类 cs.AI

Comments 34 pages, 15 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.01081 2025-04-10 cs.CV cs.CL eess.IV 57%

ShieldGemma 2: Robust and Tractable Image Content Moderation

Wenjun Zeng, Dana Kurniawan, Ryan Mullins, Yuchi Liu, Tamoghna Saha, Dirichi Ike-Njoku, Jindong Gu, Yiwen Song, Cai Xu, Jingjing Zhou, Aparna Joshi, Shravan Dheep, Mani Malek, Hamid Palangi, Joon Baek, Rick Pereira, Karthik Narasimhan

机构 * Google LLC(谷歌公司)

专题命中 安全训练 :safety(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.00441 2025-04-04 cs.CR cs.AI 57%

No Free Lunch with Guardrails

Divyanshu Kumar, Nitin Aravind Birur, Tanay Baswa, Sahil Agarwal, Prashanth Harshangi

机构 * Enkrypt AI(恩克莱普特人工智能公司)

专题命中 安全训练 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.02141 2025-04-04 cs.SE cs.AI 57%

On Simulation-Guided LLM-based Code Generation for Safe Autonomous Driving Software

Ali Nouri, Johan Andersson, Kailash De Jesus Hornig, Zhennan Fei, Emil Knabe, Hakan Sivencrona, Beatriz Cabrero-Daniel, Christian Berger

机构 * Volvo Cars(沃尔沃汽车) Chalmers University of Technology(查尔姆斯理工大学) University of Gothenburg(哥德堡大学)

专题命中 安全训练 :safety(abstract);分类 cs.AI

Comments Accepted in the 29th International Conference on Evaluation and Assessment in Software Engineering (EASE)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.01719 2025-04-04 cs.LG cs.RO 57%

Beyond Non-Expert Demonstrations: Outcome-Driven Action Constraint for Offline Reinforcement Learning

Ke Jiang, Wen Jiang, Yao Li, Xiaoyang Tan

机构 * College of Computer Science and Technology, Nanjing University of Aeronautics and Astronautics(南京航空航天大学计算机科学与技术学院) School of Computer and Information Technology, Shanxi University(山西大学计算机与信息技术学院)

专题命中 安全训练 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.18098 2025-04-02 cs.CV cs.LG 57%

Disentangling Safe and Unsafe Corruptions via Anisotropy and Locality

Ramchandran Muthukumar, Ambar Pal, Jeremias Sulam, Rene Vidal

机构 * Johns Hopkins University(约翰斯·霍普金斯大学) Amazon Web Services(亚马逊云科技) University of Pennsylvania(宾夕法尼亚大学)

专题命中 安全训练 :alignment(abstract);分类 cs.LG

Comments Published at IEEE/CVF Conference on Computer Vision and Pattern Recognition 2025. Updated Acknowledgements

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.13727 2025-04-02 cs.MA cs.AI 57%

Scalable Safe Multi-Agent Reinforcement Learning for Multi-Agent System

Haikuo Du, Fandi Gou, Yunze Cai

专题命中 安全训练 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.16871 2025-04-01 cs.MA cs.LG stat.ML 57%

Conformal Off-Policy Prediction for Multi-Agent Systems

Tom Kuipers, Renukanandan Tumu, Shuo Yang, Milad Kazemi, Rahul Mangharam, Nicola Paoletti

机构 * University of Pennsylvania(宾夕法尼亚大学)

专题命中 安全训练 :safety(abstract);分类 cs.LG

Comments Accepted for publication in the 63rd IEEE Conference on Decision and Control (CDC) 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.21141 2025-03-31 cs.RO cs.LG cs.MA 57%

Safe Human Robot Navigation in Warehouse Scenario

Seth Farrell, Chenghao Li, Hongzhan Yu, Ryo Yoshimitsu, Sicun Gao, Henrik I. Christensen

机构 * UCSD(加利福尼亚大学圣迭戈分校) IHI Japan(IHI日本公司)

专题命中 安全训练 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.21949 2025-03-31 cs.LG 57%

Reward Design for Reinforcement Learning Agents

Rati Devidze

专题命中 安全训练 :alignment(abstract);分类 cs.LG

Comments Doctoral thesis

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.20194 2025-03-27 cs.CL 57%

GAPO: Learning Preferential Prompt through Generative Adversarial Policy Optimization

Zhouhong Gu, Xingzhou Chen, Xiaoran Shi, Tao Wang, Suhang Zheng, Tianyu Li, Hongwei Feng, Yanghua Xiao

机构 * Fudan University(复旦大学) Alibaba Group(阿里巴巴集团)

专题命中 安全训练 :DPO(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.19418 2025-03-26 cs.LG 57%

Multi-Agent Deep Reinforcement Learning for Safe Autonomous Driving with RICS-Assisted MEC

Xueyao Zhang, Bo Yang, Xuelin Cao, Zhiwen Yu, George C. Alexandropoulos, Yan Zhang, Merouane Debbah, Chau Yuen

机构 * Northwestern Polytechnical University(西北工业大学) Harbin Engineering University(哈尔滨工程大学) Xidian University(西安电子科技大学) National and Kapodistrian University of Athens(雅典国立与卡波迪斯特里亚大学) University of Oslo(奥斯陆大学) Khalifa University(哈利法大学) Nanyang Technological University(南洋理工大学)

专题命中 安全训练 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.17194 2025-03-24 cs.LG 57%

Curriculum RL meets Monte Carlo Planning: Optimization of a Real World Container Management Problem

Abhijeet Pendyala, Tobias Glasmachers

机构 * Ruhr-University Bochum(波鸿鲁尔大学)

专题命中 安全训练 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.14563 2025-03-21 cs.SE cs.AI 57%

Workflow for Safe-AI

Suzana Veljanovska, Hans Dermot Doran

专题命中 安全训练 :safety(abstract);分类 cs.AI

Comments Embedded World Conference, Nuremberg, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.00959 2025-03-19 cs.CY 57%

IGGA: A Dataset of Industrial Guidelines and Policy Statements for Generative AIs

Junfeng Jiao, Saleh Afroogh, Kevin Chen, David Atkinson, Amit Dhurandhar

专题命中 安全训练 :trustworthy(abstract);分类 cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.03640 2025-03-17 cs.RO cs.LG cs.MA math.OC 57%

Discrete GCBF Proximal Policy Optimization for Multi-agent Safe Optimal Control

Songyuan Zhang, Oswin So, Mitchell Black, Chuchu Fan

专题命中 安全训练 :safety(abstract);分类 cs.LG

Comments 31 pages, 15 figures; Accepted by the thirteenth International Conference on Learning Representations (ICLR 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.06550 2025-03-11 cs.CL 57%

BingoGuard: LLM Content Moderation Tools with Risk Levels

Fan Yin, Philippe Laban, Xiangyu Peng, Yilun Zhou, Yixin Mao, Vaibhav Vats, Linnea Ross, Divyansh Agarwal, Caiming Xiong, Chien-Sheng Wu

专题命中 安全训练 :safety(abstract);分类 cs.CL

Comments 10 pages, 4 figures, 4 tables. ICLR 2025 poster

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.03911 2025-03-07 cs.RO cs.LG cs.SY eess.SY 57%

Safe LLM-Controlled Robots with Formal Guarantees via Reachability Analysis

Ahmad Hafez, Alireza Naderi Akhormeh, Amr Hegazy, Amr Alanwar

机构 * Technical University of Munich(慕尼黑工业大学) German University in Cairo(开罗德国大学)

专题命中 安全训练 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.01334 2025-03-07 cs.RO cs.AI cs.CV cs.HC 57%

A Backbone for Long-Horizon Robot Task Understanding

Xiaoshuai Chen, Wei Chen, Dongmyoung Lee, Yukun Ge, Nicolas Rojas, Petar Kormushev

机构 * Dyson School of Design Engineering, Imperial College London(帝国理工学院戴森设计工程学院) The AI Institute(人工智能研究所)

专题命中 安全训练 :alignment(abstract);分类 cs.AI

Comments 8 pages, 8 figures. This work has been published by IEEE Robotics and Automation Letters (RA-L)

Journal ref IEEE Robotics and Automation Letters, Volume: 10, 2025, 2048 - 2055

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.02690 2025-03-05 cs.CE cs.LG physics.ao-ph 57%

Generative Modeling of Microweather Wind Velocities for Urban Air Mobility

Tristan A. Shah, Michael C. Stanley, James E. Warner

机构 * NASA Langley Research Center(美国国家航空航天局兰利研究中心) Analytical Mechanics Associates(分析力学协会)

专题命中 安全训练 :safety(abstract);分类 cs.LG

Comments 17 pages, 13 figures, published in 2025 IEEE Aerospace Conference proceedings

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.03415 2025-03-05 cs.CL 57%

Surgical, Cheap, and Flexible: Mitigating False Refusal in Language Models via Single Vector Ablation

Xinpeng Wang, Chengzhi Hu, Paul Röttger, Barbara Plank

机构 * LMU Munich(慕尼黑大学) Bocconi University(博科尼大学)

专题命中 安全训练 :safety(abstract);分类 cs.CL

Comments ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏