arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 9434 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 9434 篇

2111.04972 2021-11-10 cs.LG 57%

Risk Sensitive Model-Based Reinforcement Learning using Uncertainty Guided Planning

Stefan Radic Webster, Peter Flach

专题命中 安全评测 :trustworthy(abstract);分类 cs.LG

Comments Safe RL Workshop NeurIPS 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.01743 2021-11-03 cs.LG stat.ML 57%

Designing Inherently Interpretable Machine Learning Models

Agus Sudjianto, Aijun Zhang

专题命中 安全评测 :safety(abstract);分类 cs.LG

Comments arXiv admin note: text overlap with arXiv:2011.04041

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.00961 2021-11-03 astro-ph.GA cs.CV cs.LG 57%

Robustness of deep learning algorithms in astronomy -- galaxy morphology studies

A. Ćiprijanović, D. Kafkes, G. N. Perdue, K. Pedro, G. Snyder, F. J. Sánchez, S. Madireddy, S. M. Wild, B. Nord

专题命中 安全评测 :trustworthy(abstract);分类 cs.LG

Comments Accepted in: Fourth Workshop on Machine Learning and the Physical Sciences (35th Conference on Neural Information Processing Systems; NeurIPS2021); final version

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.00881 2021-11-02 cs.LG 57%

Artificial Intelligence in the Low-Level Realm -- A Survey

Vahid Mohammadi Safarzadeh, Hamed Ghasr Loghmani

专题命中 安全评测 :trustworthy(abstract);分类 cs.LG

Comments 7 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2107.03022 2021-11-02 cs.LG 57%

Reconstructing Test Labels from Noisy Loss Functions

Abhinav Aggarwal, Shiva Prasad Kasiviswanathan, Zekun Xu, Oluwaseyi Feyisetan, Nathanael Teissier

专题命中 安全评测 :trustworthy(abstract);分类 cs.LG

Comments Accepted at NeurIPS 2021 Workshop on Privacy in Machine Learning (PriML)

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.15036 2021-10-22 cs.LG eess.SP 57%

Automated Workers Ergonomic Risk Assessment in Manual Material Handling using sEMG Wearable Sensors and Machine Learning

Srimantha E. Mudiyanselage, Phuong H. D. Nguyen, Mohammad Sadra Rajabi, Reza Akhavian

专题命中 安全评测 :safety(abstract);分类 cs.LG

Journal ref Electronics. 2021; 10(20):2558

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.09570 2021-10-20 cs.CL 57%

A Data Bootstrapping Recipe for Low Resource Multilingual Relation Classification

Arijit Nag, Bidisha Samanta, Animesh Mukherjee, Niloy Ganguly, Soumen Chakrabarti

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.07960 2021-10-19 cs.AI cs.SE 57%

Efficient and Effective Generation of Test Cases for Pedestrian Detection -- Search-based Software Testing of Baidu Apollo in SVL

Hamid Ebadi, Mahshid Helali Moghadam, Markus Borg, Gregory Gay, Afonso Fontes, Kasper Socha

专题命中 安全评测 :safety(abstract);分类 cs.AI

Comments 8 pages, 2021 IEEE Conference on Artificial Intelligence Testing (AITest 2021)

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.07658 2021-10-18 cs.LG astro-ph.SR 57%

Predicting Solar Flares with Remote Sensing and Machine Learning

Erik Larsen

专题命中 安全评测 :safety(abstract);分类 cs.LG

Comments 16 pages, 10 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.03022 2021-10-08 cs.LG stat.ML 57%

Tribuo: Machine Learning with Provenance in Java

Adam Pocock

专题命中 安全评测 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.03120 2021-10-08 cs.LG 57%

Assurance Monitoring of Learning Enabled Cyber-Physical Systems Using Inductive Conformal Prediction based on Distance Learning

Dimitrios Boursinos, Xenofon Koutsoukos

专题命中 安全评测 :trustworthy(abstract);分类 cs.LG

Comments Published in Artificial Intelligence for Engineering Design, Analysis and Manufacturing. arXiv admin note: text overlap with arXiv:2001.05014

Journal ref Artificial Intelligence for Engineering Design, Analysis and Manufacturing, 35(2), 2021, 251-264

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.00086 2021-10-04 cs.LG 57%

On the Trustworthiness of Tree Ensemble Explainability Methods

Angeline Yasodhara, Azin Asgarian, Diego Huang, Parinaz Sobhani

专题命中 安全评测 :trustworthy(abstract);分类 cs.LG

Journal ref International Cross-Domain Conference for Machine Learning and Knowledge Extraction 2021 Aug 17 (pp. 293-308). Springer, Cham

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.15207 2021-10-01 cs.CV cs.CL cs.RO 57%

Language-Aligned Waypoint (LAW) Supervision for Vision-and-Language Navigation in Continuous Environments

Sonia Raychaudhuri, Saim Wani, Shivansh Patel, Unnat Jain, Angel X. Chang

专题命中 安全评测 :alignment(abstract);分类 cs.CL

Comments EMNLP 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.14956 2021-10-01 eess.IV cs.CV cs.LG 57%

Comparative Validation of Machine Learning Algorithms for Surgical Workflow and Skill Analysis with the HeiChole Benchmark

Martin Wagner, Beat-Peter Müller-Stich, Anna Kisilenko, Duc Tran, Patrick Heger, Lars Mündermann, David M Lubotsky, Benjamin Müller, Tornike Davitashvili, Manuela Capek, Annika Reinke, Tong Yu, Armine Vardazaryan, Chinedu Innocent Nwoye, Nicolas Padoy, Xinyang Liu, Eung-Joo Lee, Constantin Disch, Hans Meine, Tong Xia, Fucang Jia, Satoshi Kondo, Wolfgang Reiter, Yueming Jin, Yonghao Long, Meirui Jiang, Qi Dou, Pheng Ann Heng, Isabell Twick, Kadir Kirtac, Enes Hosgor, Jon Lindström Bolmgren, Michael Stenzel, Björn von Siemens, Hannes G. Kenngott, Felix Nickel, Moritz von Frankenberg, Franziska Mathis-Ullrich, Lena Maier-Hein, Stefanie Speidel, Sebastian Bodenstedt

专题命中 安全评测 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.12480 2021-09-28 cs.HC cs.AI 57%

Explainability Pitfalls: Beyond Dark Patterns in Explainable AI

Upol Ehsan, Mark O. Riedl

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.07905 2021-09-17 cs.CY 57%

Risk Management of AI/ML Software as a Medical Device (SaMD): On ISO 14971 and Related Standards and Guidances

Stephen G. Odaibo

专题命中 安全评测 :safety(abstract);分类 cs.CY

Comments 25 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2102.07617 2021-09-15 cs.AI 57%

On the Philosophical, Cognitive and Mathematical Foundations of Symbiotic Autonomous Systems (SAS)

Yingxu Wang, Fakhri Karray, Sam Kwong, Konstantinos N. Plataniotis, Henry Leung, Ming Hou, Edward Tunstel, Imre J. Rudas, Ljiljana Trajkovic, Okyay Kaynak, Janusz Kacprzyk, Mengchu Zhou, Michael H. Smith, Philip Chen, Shushma Patel

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

Comments Accepted by Phil. Trans. Royal Society (A): Math, Phys & Engg Sci., 379(219x), 2021, Oxford, UK

Journal ref Phil. Trans. Royal Society (A): Math, Phys & Engg Sci., 379(219x), 2021, Oxford, UK

详情

展开后加载摘要…

URL PDF HTML 收藏
2108.13930 2021-09-03 cs.CR cs.LG 57%

EG-Booster: Explanation-Guided Booster of ML Evasion Attacks

Abderrahmen Amich, Birhanu Eshete

专题命中 安全评测 :trustworthy(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2108.10168 2021-08-24 cs.AI 57%

CGEMs: A Metric Model for Automatic Code Generation using GPT-3

Aishwarya Narasimhan, Krishna Prasad Agara Venkatesha Rao, Veena M B

专题命中 安全评测 :safety(abstract);分类 cs.AI

Comments 11 pages, 6 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2108.08759 2021-08-20 cs.CL 57%

DESYR: Definition and Syntactic Representation Based Claim Detection on the Web

Megha Sundriyal, Parantak Singh, Md Shad Akhtar, Shubhashis Sengupta, Tanmoy Chakraborty

专题命中 安全评测 :alignment(abstract);分类 cs.CL

Comments 10 pages, Accepted at CIKM 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2009.14645 2021-08-11 cs.CE cs.LG cs.SY eess.SY 57%

Computational framework for real-time diagnostics and prognostics of aircraft actuation systems

Pier Carlo Berri, Matteo D. L. Dalla Vedova, Laura Mainini

专题命中 安全评测 :safety(abstract);分类 cs.LG

Comments 57 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2104.06982 2021-07-29 cs.AI 57%

To Trust or Not to Trust a Regressor: Estimating and Explaining Trustworthiness of Regression Predictions

Kim de Bie, Ana Lucic, Hinda Haned

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

Comments Accepted to ICML 2021 Workshop on Human in the Loop Learning (HILL)

详情

展开后加载摘要…

URL PDF HTML 收藏
2106.09885 2021-07-23 eess.AS cs.AI 57%

An Improved Single Step Non-autoregressive Transformer for Automatic Speech Recognition

Ruchao Fan, Wei Chu, Peng Chang, Jing Xiao, Abeer Alwan

专题命中 安全评测 :alignment(abstract);分类 cs.AI

Comments Accepted to Interspeech2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2107.03374 2021-07-15 cs.LG 57%

Evaluating Large Language Models Trained on Code

Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, Alex Ray, Raul Puri, Gretchen Krueger, Michael Petrov, Heidy Khlaaf, Girish Sastry, Pamela Mishkin, Brooke Chan, Scott Gray, Nick Ryder, Mikhail Pavlov, Alethea Power, Lukasz Kaiser, Mohammad Bavarian, Clemens Winter, Philippe Tillet, Felipe Petroski Such, Dave Cummings, Matthias Plappert, Fotios Chantzis, Elizabeth Barnes, Ariel Herbert-Voss, William Hebgen Guss, Alex Nichol, Alex Paino, Nikolas Tezak, Jie Tang, Igor Babuschkin, Suchir Balaji, Shantanu Jain, William Saunders, Christopher Hesse, Andrew N. Carr, Jan Leike, Josh Achiam, Vedant Misra, Evan Morikawa, Alec Radford, Matthew Knight, Miles Brundage, Mira Murati, Katie Mayer, Peter Welinder, Bob McGrew, Dario Amodei, Sam McCandlish, Ilya Sutskever, Wojciech Zaremba

专题命中 安全评测 :safety(abstract);分类 cs.LG

Comments corrected typos, added references, added authors, added acknowledgements

详情

展开后加载摘要…

URL PDF HTML 收藏
2107.06098 2021-07-14 cs.LG cs.CV 57%

Using Causal Analysis for Conceptual Deep Learning Explanation

Sumedha Singla, Stephen Wallace, Sofia Triantafillou, Kayhan Batmanghelich

专题命中 安全评测 :trustworthy(abstract);分类 cs.LG

Comments 10 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
1902.07247 2021-06-22 cs.LG stat.ML 57%

Fast Neural Network Verification via Shadow Prices

Vicenc Rubies-Royo, Roberto Calandra, Dusan M. Stipanovic, Claire Tomlin

专题命中 安全评测 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2106.09223 2021-06-18 cs.LG cs.CV 57%

Evaluating the Robustness of Bayesian Neural Networks Against Different Types of Attacks

Yutian Pang, Sheng Cheng, Jueming Hu, Yongming Liu

专题命中 安全评测 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2106.01623 2021-06-04 cs.CL 57%

Few-shot Knowledge Graph-to-Text Generation with Pretrained Language Models

Junyi Li, Tianyi Tang, Wayne Xin Zhao, Zhicheng Wei, Nicholas Jing Yuan, Ji-Rong Wen

专题命中 安全评测 :alignment(abstract);分类 cs.CL

Comments Accepted to ACL 2021 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2103.02630 2021-06-02 cs.LG 57%

Hypothesis Testing for Class-Conditional Label Noise

Rafael Poyiadzi, Weisong Yang, Niall Twomey, Raul Santos-Rodriguez

专题命中 安全评测 :trustworthy(abstract);分类 cs.LG

Comments 15 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2105.12700 2021-05-28 eess.IV cs.CV cs.LG cs.MM 57%

Towards Transparent Application of Machine Learning in Video Processing

Luka Murn, Marc Gorriz Blanch, Maria Santamaria, Fiona Rivera, Marta Mrak

专题命中 安全评测 :trustworthy(abstract);分类 cs.LG

Comments International Broadcasting Convention, 11-14 Sep 2020, Amsterdam, Netherlands (Technical Paper section, Virtual)

详情

展开后加载摘要…

URL PDF HTML 收藏