arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 3309 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全训练 3309 篇

2210.01241 2023-12-05 cs.CL cs.LG 62%

Is Reinforcement Learning (Not) for Natural Language Processing: Benchmarks, Baselines, and Building Blocks for Natural Language Policy Optimization

Rajkumar Ramamurthy, Prithviraj Ammanabrolu, Kianté Brantley, Jack Hessel, Rafet Sifa, Christian Bauckhage, Hannaneh Hajishirzi, Yejin Choi

专题命中 安全训练 :alignment(abstract);分类 cs.CL、cs.LG

Comments In Proceedings of ICLR 2023. Code found at https://github.com/allenai/rl4lms and Project website at https://rl4lms.apps.allenai.org/

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.17769 2023-12-05 cs.CL cs.AI 62%

Social Contract AI: Aligning AI Assistants with Implicit Group Norms

Jan-Philipp Fränken, Sam Kwok, Peixuan Ye, Kanishk Gandhi, Dilip Arumugam, Jared Moore, Alex Tamkin, Tobias Gerstenberg, Noah D. Goodman

专题命中 安全训练 :alignment(abstract);分类 cs.CL、cs.AI

Comments SoLaR NeurIPS 2023 Workshop (https://solar-neurips.github.io/)

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.08592 2023-12-01 cs.SE cs.AI cs.CL 62%

AART: AI-Assisted Red-Teaming with Diverse Data Generation for New LLM-powered Applications

Bhaktipriya Radharapu, Kevin Robinson, Lora Aroyo, Preethi Lahoti

专题命中 安全训练 :safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.10863 2023-11-23 cs.RO cs.AI cs.LG 62%

Verified Compositional Neuro-Symbolic Control for Stochastic Systems with Temporal Logic Tasks

Jun Wang, Haojun Chen, Zihe Sun, Yiannis Kantaros

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG

Comments arXiv admin note: substantial text overlap with arXiv:2209.06130

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.03033 2023-11-07 cs.LG cs.AI 62%

Beyond Words: A Mathematical Framework for Interpreting Large Language Models

Javier González, Aditya V. Nori

专题命中 安全训练 :alignment(abstract);分类 cs.AI、cs.LG

Comments 4 figures, 18 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.00063 2023-11-02 cs.RO cs.AI cs.LG cs.MA cs.SY eess.SY 62%

Safe multi-agent motion planning under uncertainty for drones using filtered reinforcement learning

Sleiman Safaoui, Abraham P. Vinod, Ankush Chakrabarty, Rien Quirynen, Nobuyuki Yoshikawa, Stefano Di Cairano

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.20287 2023-11-01 cs.LG cs.AI 62%

Sample-Efficient and Safe Deep Reinforcement Learning via Reset Deep Ensemble Agents

Woojun Kim, Yongjae Shin, Jongeui Park, Youngchul Sung

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG

Comments NeurIPS 2023 camera-ready

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.14788 2023-10-24 cs.LG cs.AI cs.SY eess.SY 62%

Specialized Deep Residual Policy Safe Reinforcement Learning-Based Controller for Complex and Continuous State-Action Spaces

Ammar N. Abbas, Georgios C. Chasparis, John D. Kelleher

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.10501 2023-10-17 cs.CL cs.AI 62%

NeMo Guardrails: A Toolkit for Controllable and Safe LLM Applications with Programmable Rails

Traian Rebedea, Razvan Dinu, Makesh Sreedhar, Christopher Parisien, Jonathan Cohen

专题命中 安全训练 :alignment(abstract);分类 cs.CL、cs.AI

Comments Accepted at EMNLP 2023 - Demo track

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.08595 2023-10-17 cs.RO cs.AI cs.LG 62%

Deep Reinforcement Learning for Autonomous Vehicle Intersection Navigation

Badr Ben Elallid, Hamza El Alaoui, Nabil Benamar

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG

Comments Accepted for publication in the 2023 International Conference on Innovation and Intelligence for Informatics, Computing, and Technologies (3ICT)

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.10007 2023-10-03 cs.RO cs.AI cs.LG cs.MA 62%

Multi-Agent Deep Reinforcement Learning for Cooperative and Competitive Autonomous Vehicles using AutoDRIVE Ecosystem

Tanmay Vilas Samak, Chinmay Vilas Samak, Venkat Krovi

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG

Comments Accepted as Multi-Agent Dynamic Games (MAD-Games) Workshop Paper at IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.04707 2023-09-12 cs.AI cs.LG 62%

Advantage Actor-Critic with Reasoner: Explaining the Agent's Behavior from an Exploratory Perspective

Muzhe Guo, Feixu Yu, Tian Lan, Fang Jin

专题命中 安全训练 :trustworthy(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.12270 2023-08-24 cs.LG cs.AI 62%

Language Reward Modulation for Pretraining Reinforcement Learning

Ademi Adeniji, Amber Xie, Carmelo Sferrazza, Younggyo Seo, Stephen James, Pieter Abbeel

专题命中 安全训练 :alignment(abstract);分类 cs.AI、cs.LG

Comments Code available at https://github.com/ademiadeniji/lamp

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.10130 2023-08-22 econ.GN cs.AI cs.CY q-fin.EC 62%

GPTs are GPTs: An Early Look at the Labor Market Impact Potential of Large Language Models

Tyna Eloundou, Sam Manning, Pamela Mishkin, Daniel Rock

专题命中 安全训练 :alignment(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.03907 2023-08-09 cs.HC cs.CY cs.LG 62%

Advancements In Crowd-Monitoring System: A Comprehensive Analysis of Systematic Approaches and Automation Algorithms: State-of-The-Art

Mohammed Ameen, Richard Stone

专题命中 安全训练 :safety(abstract);分类 cs.CY、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.07813 2023-08-08 cs.SE cs.AI cs.LG 62%

A Search-Based Testing Approach for Deep Reinforcement Learning Agents

Amirhossein Zolfagharian, Manel Abdellatif, Lionel Briand, Mojtaba Bagherzadeh, Ramesh S

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG

Journal ref in IEEE Transactions on Software Engineering, vol. 49, no. 7, pp. 3715-3735, July 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.13944 2023-06-27 cs.LG cs.AI 62%

Safe Reinforcement Learning with Dead-Ends Avoidance and Recovery

Xiao Zhang, Hai Zhang, Hongtu Zhou, Chang Huang, Di Zhang, Chen Ye, Junqiao Zhao

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG

Comments 8 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.02670 2023-06-21 cs.LG cs.AI cs.CR cs.RO 62%

Robust Adversarial Attacks Detection based on Explainable Deep Reinforcement Learning For UAV Guidance and Planning

Thomas Hickling, Nabil Aouf, Phillippa Spencer

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG

Comments 13 pages, 16 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.09055 2023-06-16 cs.RO cs.AI cs.LG 62%

Predictive Maneuver Planning with Deep Reinforcement Learning (PMP-DRL) for comfortable and safe autonomous driving

Jayabrata Chowdhury, Vishruth Veerendranath, Suresh Sundaram, Narasimhan Sundararajan

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.02337 2023-05-15 cs.CY cs.AI 62%

Regulating ChatGPT and other Large Generative AI Models

Philipp Hacker, Andreas Engel, Marco Mauer

专题命中 安全训练 :trustworthy(abstract);分类 cs.AI、cs.CY

Comments FAccT '23, June 12-15, 2023, Chicago, IL, USA

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.13191 2023-04-27 cs.AI cs.CL 62%

Towards Explainable and Safe Conversational Agents for Mental Health: A Survey

Surjodeep Sarkar, Manas Gaur, L. Chen, Muskan Garg, Biplav Srivastava, Bhaktee Dongaonkar

专题命中 安全训练 :safety(abstract);分类 cs.CL、cs.AI

Comments 10 pages, 3 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.01728 2023-04-25 cs.LG cs.AI cs.RO 62%

Guarded Policy Optimization with Imperfect Online Demonstrations

Zhenghai Xue, Zhenghao Peng, Quanyi Li, Zhihan Liu, Bolei Zhou

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG

Comments Accepted at ICLR 2023 (top 25%)

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.12558 2023-04-24 cs.LG cs.AI 62%

Wasserstein Auto-encoded MDPs: Formal Verification of Efficiently Distilled RL Policies with Many-sided Guarantees

Florent Delgrange, Ann Nowé, Guillermo A. Pérez

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG

Comments ICLR 2023, 10 pages main text, 14 pages appendix (excluding references)

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.09448 2023-04-20 cs.LG cs.CL cs.CV 62%

EC^2: Emergent Communication for Embodied Control

Yao Mu, Shunyu Yao, Mingyu Ding, Ping Luo, Chuang Gan

专题命中 安全训练 :alignment(abstract);分类 cs.CL、cs.LG

Comments Published in CVPR2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.06281 2023-04-14 cs.LG cs.AI cs.MA cs.RO 62%

Model-based Dynamic Shielding for Safe and Efficient Multi-Agent Reinforcement Learning

Wenli Xiao, Yiwei Lyu, John Dolan

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG

Comments Accepted in AAMAS 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.03081 2023-04-07 cs.LG cs.AI 62%

Safe MDP Planning by Learning Temporal Patterns of Undesirable Trajectories and Averting Negative Side Effects

Siow Meng Low, Akshat Kumar, Scott Sanner

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.15452 2023-03-21 cs.LG cs.AI cs.SY eess.SY stat.ML 62%

Safe Exploration Method for Reinforcement Learning under Existence of Disturbance

Yoshihiro Okawa, Tomotake Sasaki, Hitoshi Yanami, Toru Namerikawa

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG

Comments Accepted by the European Conference on Machine Learning and Principles and Practice of Knowledge Discovery in Databases (ECMLPKDD) 2022. The Version of Record is available at https://doi.org/10.1007/978-3-031-26412-2_9

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.08187 2023-03-16 cs.LG cs.AI 62%

Vehicle lateral control using Machine Learning for automated vehicle guidance

Akash Fogla, Kanish Kumar, Sunnay Saurav, Bishnu ramanujan

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.10416 2023-03-15 cs.LG cs.CY eess.SP 62%

CycleSense: Detecting Near Miss Incidents in Bicycle Traffic from Mobile Motion Sensors

Ahmet-Serdar Karakaya, Thomas Ritter, Felix Biessmann, David Bermbach

专题命中 安全训练 :safety(abstract);分类 cs.CY、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.03226 2023-03-07 cs.AI cs.LG 62%

Safe Reinforcement Learning via Probabilistic Logic Shields

Wen-Chi Yang, Giuseppe Marra, Gavin Rens, Luc De Raedt

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏