arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 3309 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全训练 3309 篇

2102.13045 2021-02-26 cs.LG cs.AI 62%

Iterative Bounding MDPs: Learning Interpretable Policies via Non-Interpretable Methods

Nicholay Topin, Stephanie Milani, Fei Fang, Manuela Veloso

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2102.12432 2021-02-25 cs.RO cs.AI cs.LG cs.SY eess.SY 62%

Deep Reinforcement Learning for Safe Landing Site Selection with Concurrent Consideration of Divert Maneuvers

Keidai Iiyama, Kento Tomita, Bhavi A. Jagatia, Tatsuwaki Nakagawa, Koki Ho

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG

Comments 25 pages, 14 figures, This paper is an updated version of Paper AAS 20-583 presented at the AAS/AIAA Astrodynamics Specialist Conference, Online

详情

展开后加载摘要…

URL PDF HTML 收藏
2006.12136 2021-01-22 cs.LG cs.AI cs.RO 62%

Safe Reinforcement Learning via Curriculum Induction

Matteo Turchetta, Andrey Kolobov, Shital Shah, Andreas Krause, Alekh Agarwal

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2008.10148 2020-08-25 cs.HC cs.AI cs.LG cs.SY eess.SY 62%

Drive Safe: Cognitive-Behavioral Mining for Intelligent Transportation Cyber-Physical System

Md. Shirajum Munir, Sarder Fakhrul Abedin, Ki Tae Kim, Do Hyeon Kim, Md. Golam Rabiul Alam, Choong Seon Hong

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG

Comments Submitted to IEEE Transactions on Intelligent Transportation Systems, Special Issue on Technologies for risk mitigation and support of impaired drivers

详情

展开后加载摘要…

URL PDF HTML 收藏
2008.06696 2020-08-19 cs.AI cs.LG cs.RO 62%

Autonomous Braking and Throttle System: A Deep Reinforcement Learning Approach for Naturalistic Driving

Varshit S. Dubey, Ruhshad Kasad, Karan Agrawal

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2008.06626 2020-08-18 cs.LG cs.AI cs.RO 62%

Safe Reinforcement Learning in Constrained Markov Decision Processes

Akifumi Wachi, Yanan Sui

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG

Comments 10 pages, 6 figures, Accepted to ICML2020

详情

展开后加载摘要…

URL PDF HTML 收藏
1902.00089 2020-07-14 cs.LG cs.AI cs.RO stat.ML 62%

Safe, Efficient, and Comfortable Velocity Control based on Reinforcement Learning for Autonomous Driving

Meixin Zhu, Yinhai Wang, Ziyuan Pu, Jingyun Hu, Xuesong Wang, Ruimin Ke

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG

Comments Under the first-round revision for transportation research part c

Journal ref Transportation Research Part C: Emerging Technologies 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2005.12069 2020-05-26 cs.LG cs.AI stat.ML 62%

Policy Entropy for Out-of-Distribution Classification

Andreas Sedlmeier, Robert Müller, Steffen Illium, Claudia Linnhoff-Popien

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2002.02667 2020-05-22 cs.LG cs.AI cs.RO eess.SP 62%

Automated Lane Change Strategy using Proximal Policy Optimization-based Deep Reinforcement Learning

Fei Ye, Xuxin Cheng, Pin Wang, Ching-Yao Chan, Jiucai Zhang

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2004.08648 2020-04-21 cs.LG cs.AI cs.RO stat.ML 62%

Modeling Survival in model-based Reinforcement Learning

Saeed Moazami, Peggy Doerschuk

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2004.03237 2020-04-08 cs.LG cs.AI cs.NE 62%

How Do You Act? An Empirical Study to Understand Behavior of Deep Reinforcement Learning Agents

Richard Meyes, Moritz Schneider, Tobias Meisen

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG

Comments 16 pages, currently under review for publication for the ECMLPKDD 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2004.00716 2020-04-03 cs.RO cs.AI cs.LG 62%

Constrained-Space Optimization and Reinforcement Learning for Complex Tasks

Ya-Yen Tsai, Bo Xiao, Edward Johns, Guang-Zhong Yang

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG

Comments Accepted for publication in RA-Letters and at ICRA 2020

Journal ref IEEE Robotics and Automation Letters, 5(2) (2020) 682-689

详情

展开后加载摘要…

URL PDF HTML 收藏
2002.12156 2020-03-24 cs.LG cs.AI cs.LO cs.SY eess.SY stat.ML 62%

Cautious Reinforcement Learning with Logical Constraints

Mohammadhosein Hasanbeig, Alessandro Abate, Daniel Kroening

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG

Comments Accepted to AAMAS 2020. arXiv admin note: text overlap with arXiv:1902.00778

详情

展开后加载摘要…

URL PDF HTML 收藏
1911.11699 2020-02-12 cs.RO cs.AI cs.LG cs.MA 62%

Multi-Vehicle Mixed-Reality Reinforcement Learning for Autonomous Multi-Lane Driving

Rupert Mitchell, Jenny Fletcher, Jacopo Panerati, Amanda Prorok

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG

Comments 10 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
1910.02390 2019-11-25 cs.LG cs.CY stat.ML 62%

Migration through Machine Learning Lens -- Predicting Sexual and Reproductive Health Vulnerability of Young Migrants

Amber Nigam, Pragati Jaiswal, Uma Girkar, Teertha Arora, Leo A. Celi

专题命中 安全训练 :safety(abstract);分类 cs.CY、cs.LG

Comments Accepted for Machine Learning for Health (ML4H) at NeurIPS 2019 - Extended Abstract

详情

展开后加载摘要…

URL PDF HTML 收藏
1910.13726 2019-10-31 cs.LG cs.AI cs.RO stat.ML 62%

Safe Exploration for Interactive Machine Learning

Matteo Turchetta, Felix Berkenkamp, Andreas Krause

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG

Comments Accepted at NeurIPS 2019

详情

展开后加载摘要…

URL PDF HTML 收藏
1906.12189 2019-07-02 eess.SY cs.AI cs.LG cs.SY 62%

Learning-based Model Predictive Control for Safe Exploration and Reinforcement Learning

Torsten Koller, Felix Berkenkamp, Matteo Turchetta, Joschka Boedecker, Andreas Krause

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG

Comments 14 pages, 7 figures. arXiv admin note: substantial text overlap with arXiv:1803.08287

详情

展开后加载摘要…

URL PDF HTML 收藏
1904.07705 2019-07-02 cs.RO cs.AI cs.LG 62%

Multi-Objective Autonomous Braking System using Naturalistic Dataset

Rafael Vasquez, Bilal Farooq

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG

Comments Accepted in the proceedings of IEEE ITSC2019

详情

展开后加载摘要…

URL PDF HTML 收藏
1905.04127 2019-05-13 cs.LG cs.AI 62%

Design of Artificial Intelligence Agents for Games using Deep Reinforcement Learning

Andrei Claudiu Roibu

专题命中 安全训练 :alignment(abstract);分类 cs.AI、cs.LG

Comments Dissertation submitted to the University of Sheffield in partial fulfilment of the requirements for the degree of Master of Engineering. 98 pages, 21 Tables, 58 Figures

详情

展开后加载摘要…

URL PDF HTML 收藏
1905.01303 2019-05-07 cs.LG cs.AI cs.MA stat.ML 62%

Autonomous Air Traffic Controller: A Deep Multi-Agent Reinforcement Learning Approach

Marc Brittain, Peng Wei

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG

Comments 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
1809.00549 2019-03-26 cs.CL cs.AI 62%

Emergence of Communication in an Interactive World with Consistent Speakers

Ben Bogin, Mor Geva, Jonathan Berant

专题命中 安全训练 :alignment(abstract);分类 cs.CL、cs.AI

Comments Emergent Communication Workshop @ NeurIPS 2018

详情

展开后加载摘要…

URL PDF HTML 收藏
1810.08700 2019-03-04 cs.RO cs.AI cs.LG 62%

Safe Reinforcement Learning with Model Uncertainty Estimates

Björn Lütjens, Michael Everett, Jonathan P. How

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG

Comments ICRA 2019; Presented at IROS 2018 Workshop on Machine Learning in Robot Motion Planning

详情

展开后加载摘要…

URL PDF HTML 收藏
1812.09724 2018-12-27 cs.LG cs.AI cs.HC 62%

Parallelized Interactive Machine Learning on Autonomous Vehicles

Xi Chen, Caylin Hickey

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG

Comments 6 pages, NAECON 2018 - IEEE National Aerospace and Electronics Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
1809.08925 2018-09-25 cs.LG cs.AI 62%

Constrained Exploration and Recovery from Experience Shaping

Tu-Hoa Pham, Giovanni De Magistris, Don Joven Agravante, Subhajit Chaudhury, Asim Munawar, Ryuki Tachibana

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG

Comments Code: https://github.com/IBM/constrained-rl

详情

展开后加载摘要…

URL PDF HTML 收藏
1807.00412 2018-09-12 cs.LG cs.AI cs.RO stat.ML 62%

Learning to Drive in a Day

Alex Kendall, Jeffrey Hawke, David Janz, Przemyslaw Mazur, Daniele Reda, John-Mark Allen, Vinh-Dieu Lam, Alex Bewley, Amar Shah

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG

Comments Further results and demo videos can be viewed at: https://wayve.ai/blog/l2diad

详情

展开后加载摘要…

URL PDF HTML 收藏
1805.12487 2018-07-06 cs.LG cs.AI cs.CR cs.CV stat.ML 62%

Sequential Attacks on Agents for Long-Term Adversarial Goals

Edgar Tretschk, Seong Joon Oh, Mario Fritz

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
1805.07708 2018-05-22 cs.LG cs.AI stat.ML 62%

A Lyapunov-based Approach to Safe Reinforcement Learning

Yinlam Chow, Ofir Nachum, Edgar Duenez-Guzman, Mohammad Ghavamzadeh

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
1803.07170 2018-03-22 cs.AI cs.CY 62%

Blaming humans in autonomous vehicle accidents: Shared responsibility across levels of automation

Edmond Awad, Sydney Levine, Max Kleiman-Weiner, Sohan Dsouza, Joshua B. Tenenbaum, Azim Shariff, Jean-François Bonnefon, Iyad Rahwan

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
1708.08611 2017-09-05 cs.LO cs.AI cs.LG 62%

Safe Reinforcement Learning via Shielding

Mohammed Alshiekh, Roderick Bloem, Ruediger Ehlers, Bettina Könighofer, Scott Niekum, Ufuk Topcu

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
1610.03295 2016-10-12 cs.AI cs.LG stat.ML 62%

Safe, Multi-Agent, Reinforcement Learning for Autonomous Driving

Shai Shalev-Shwartz, Shaked Shammah, Amnon Shashua

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏