arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 9434 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 9434 篇

2303.07125 2023-03-15 cs.LG cs.CV 57%

Don't PANIC: Prototypical Additive Neural Network for Interpretable Classification of Alzheimer's Disease

Tom Nuno Wolf, Sebastian Pölsterl, Christian Wachinger

专题命中 安全评测 :trustworthy(abstract);分类 cs.LG

Comments To be published in proceedings of Information Processing In Medical Imaging 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.00293 2023-03-02 cs.CL 57%

How Robust is GPT-3.5 to Predecessors? A Comprehensive Study on Language Understanding Tasks

Xuanting Chen, Junjie Ye, Can Zu, Nuo Xu, Rui Zheng, Minlong Peng, Jie Zhou, Tao Gui, Qi Zhang, Xuanjing Huang

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.14803 2023-03-01 cs.RO cs.LG 57%

Learned Risk Metric Maps for Kinodynamic Systems

Ross Allen, Wei Xiao, Daniela Rus

专题命中 安全评测 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.13115 2023-02-28 cs.AI 57%

Dual Formulation for Chance Constrained Stochastic Shortest Path with Application to Autonomous Vehicle Behavior Planning

Rashid Alyassi, Majid Khonji

专题命中 安全评测 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.12691 2023-02-27 cs.AI 57%

Reproducibility of Machine Learning: Terminology, Recommendations and Open Issues

Riccardo Albertoni, Sara Colantonio, Piotr Skrzypczyński, Jerzy Stefanowski

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.13058 2023-02-27 cs.LG cs.CR 57%

Adversarial Robustness for Tabular Data through Cost and Utility Awareness

Klim Kireev, Bogdan Kulynych, Carmela Troncoso

专题命中 安全评测 :safety(abstract);分类 cs.LG

Comments The first two authors contributed equally. To appear in the proceedings of NDSS 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.11937 2023-02-22 cs.AI 57%

Adversarial Deep Reinforcement Learning for Improving the Robustness of Multi-agent Autonomous Driving Policies

Aizaz Sharif, Dusica Marijan

专题命中 安全评测 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.08418 2023-02-17 cs.AI cs.NI 57%

Generative AI-empowered Simulation for Autonomous Driving in Vehicular Mixed Reality Metaverses

Minrui Xu, Dusit Niyato, Junlong Chen, Hongliang Zhang, Jiawen Kang, Zehui Xiong, Shiwen Mao, Zhu Han

专题命中 安全评测 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.10895 2023-02-15 cs.CV cs.AI 57%

A Comprehensive Study of Real-Time Object Detection Networks Across Multiple Domains: A Survey

Elahe Arani, Shruthi Gowda, Ratnajit Mukherjee, Omar Magdy, Senthilkumar Kathiresan, Bahram Zonooz

专题命中 安全评测 :safety(abstract);分类 cs.AI

Comments Published in Transactions on Machine Learning Research (TMLR) with Survey Certification

Journal ref Transactions on Machine Learning Research, 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.12540 2023-02-14 cs.CL 57%

EntityCS: Improving Zero-Shot Cross-lingual Transfer with Entity-Centric Code Switching

Chenxi Whitehouse, Fenia Christopoulou, Ignacio Iacobacci

专题命中 安全评测 :alignment(abstract);分类 cs.CL

Comments Findings of EMNLP 2022

Journal ref Findings of the Association for Computational Linguistics: EMNLP 2022 (6698-6714)

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.06409 2023-02-14 cs.AI cs.SE 57%

Capabilities for Better ML Engineering

Chenyang Yang, Rachel Brower-Sinning, Grace A. Lewis, Christian Kästner, Tongshuang Wu

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.05196 2023-02-13 cs.LG 57%

Two-step counterfactual generation for OOD examples

Nawid Keshtmand, Raul Santos-Rodriguez, Jonathan Lawry

专题命中 安全评测 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.10861 2023-02-13 cs.LG math.DS 57%

A novel corrective-source term approach to modeling unknown physics in aluminum extraction process

Haakon Robinson, Erlend Lundby, Adil Rasheed, Jan Tommy Gravdahl

专题命中 安全评测 :trustworthy(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.13915 2023-02-10 cs.LG cs.LO 57%

Towards Formal XAI: Formally Approximate Minimal Explanations of Neural Networks

Shahaf Bassan, Guy Katz

专题命中 安全评测 :safety(abstract);分类 cs.LG

Comments To appear in Proc. 29th Int. Conf. on Tools and Algorithms for the Construction and Analysis of Systems (TACAS)

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.04135 2023-02-09 cs.CV cs.AI 57%

Multi-Modal Evaluation Approach for Medical Image Segmentation

Seyed M. R. Modaresi, Aomar Osmani, Mohammadreza Razzazi, Abdelghani Chibani

专题命中 安全评测 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.00775 2023-02-03 cs.LG 57%

Model Monitoring and Robustness of In-Use Machine Learning Models: Quantifying Data Distribution Shifts Using Population Stability Index

Aria Khademi, Michael Hopka, Devesh Upadhyay

专题命中 安全评测 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.00750 2023-02-03 cs.CR cs.AI 57%

Developing Hands-on Labs for Source Code Vulnerability Detection with AI

Maryam Taeb

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.11324 2023-01-27 cs.LG 57%

Certified Interpretability Robustness for Class Activation Mapping

Alex Gu, Tsui-Wei Weng, Pin-Yu Chen, Sijia Liu, Luca Daniel

专题命中 安全评测 :safety(abstract);分类 cs.LG

Comments 13 pages, 5 figures. Accepted to Machine Learning for Autonomous Driving Workshop at NeurIPS 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.11112 2023-01-27 cs.SI cs.AI 57%

uHelp: intelligent volunteer search for mutual help communities

Nardine Osman, Bruno Rosell, Carles Sierra, Marco Schorlemmer, Jordi Sabater-Mir, Lissette Lemus

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.09515 2023-01-24 cs.LG cs.CV 57%

StyleGAN-T: Unlocking the Power of GANs for Fast Large-Scale Text-to-Image Synthesis

Axel Sauer, Tero Karras, Samuli Laine, Andreas Geiger, Timo Aila

专题命中 安全评测 :alignment(abstract);分类 cs.LG

Comments Project page: https://sites.google.com/view/stylegan-t/

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.08488 2023-01-23 cs.CY 57%

Towards Openness Beyond Open Access: User Journeys through 3 Open AI Collaboratives

Jennifer Ding, Christopher Akiki, Yacine Jernite, Anne Lee Steele, Temi Popo

专题命中 安全评测 :trustworthy(abstract);分类 cs.CY

Comments Presented at the 2022 NeurIPS Workshop on Broadening Research Collaborations in ML

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.11123 2022-12-22 cs.CV cs.AI cs.RO 57%

THMA: Tencent HD Map AI System for Creating HD Map Annotations

Kun Tang, Xu Cao, Zhipeng Cao, Tong Zhou, Erlong Li, Ao Liu, Shengtao Zou, Chang Liu, Shuqi Mei, Elena Sizikova, Chao Zheng

专题命中 安全评测 :safety(abstract);分类 cs.AI

Comments IAAI 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.10628 2022-12-22 cs.CR cs.LG 57%

Holistic risk assessment of inference attacks in machine learning

Yang Yang

专题命中 安全评测 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.08550 2022-12-19 cs.CL 57%

Fine-grained Czech News Article Dataset: An Interdisciplinary Approach to Trustworthiness Analysis

Matyáš Boháček, Michal Bravanský, Filip Trhlík, Václav Moravec

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL

Comments 13 pages, 3 figures; to be published at the Second Workshop on Multimodal Fact-Checking and Hate Speech Detection (DEFACTIFY 2023) at the AAAI 2023 Conference, February 14, 2023, Washington, D.C

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.08388 2022-12-19 cs.CL 57%

Homonymy Information for English WordNet

Rowan Hall Maudslay, Simone Teufel

专题命中 安全评测 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.03275 2022-12-14 cs.LG 57%

Achieving and Understanding Out-of-Distribution Generalization in Systematic Reasoning in Small-Scale Transformers

Andrew J. Nam, Mustafa Abdool, Trevor Maxfield, James L. McClelland

专题命中 安全评测 :alignment(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.09817 2022-12-08 cs.CV cs.CL 57%

Making the Most of Text Semantics to Improve Biomedical Vision--Language Processing

Benedikt Boecking, Naoto Usuyama, Shruthi Bannur, Daniel C. Castro, Anton Schwaighofer, Stephanie Hyland, Maria Wetscherek, Tristan Naumann, Aditya Nori, Javier Alvarez-Valle, Hoifung Poon, Ozan Oktay

专题命中 安全评测 :alignment(abstract);分类 cs.CL

Comments To appear in ECCV 2022. Code: https://aka.ms/biovil-code Dataset: https://aka.ms/ms-cxr Demo Notebook: https://aka.ms/biovil-demo-notebook

Journal ref Computer Vision - ECCV 2022, LNCS vol 13696, pp 1-21

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.01861 2022-12-06 cs.LG cs.LO 57%

Online Shielding for Reinforcement Learning

Bettina Könighofer, Julian Rudolf, Alexander Palmisano, Martin Tappler, Roderick Bloem

专题命中 安全评测 :safety(abstract);分类 cs.LG

Comments arXiv admin note: substantial text overlap with arXiv:2012.09539

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.01838 2022-12-06 cs.LG cs.LO 57%

Automata Learning meets Shielding

Martin Tappler, Stefan Pranger, Bettina Könighofer, Edi Muškardin, Roderick Bloem, Kim Larsen

专题命中 安全评测 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2201.10459 2022-12-06 cs.LG cs.DB 57%

FRAMED: An AutoML Approach for Structural Performance Prediction of Bicycle Frames

Lyle Regenwetter, Colin Weaver, Faez Ahmed

专题命中 安全评测 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏