arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 9400 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 9400 篇

2302.09190 2023-02-21 cs.LG cs.CY 81%

Function Composition in Trustworthy Machine Learning: Implementation Choices, Insights, and Questions

Manish Nagireddy, Moninder Singh, Samuel C. Hoffman, Evaline Ju, Karthikeyan Natesan Ramamurthy, Kush R. Varshney

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.CY、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.00646 2023-01-16 cs.SE cs.AI cs.LG cs.RO 81%

Reliability Assessment and Safety Arguments for Machine Learning Components in System Assurance

Yi Dong, Wei Huang, Vibhav Bharti, Victoria Cox, Alec Banks, Sen Wang, Xingyu Zhao, Sven Schewe, Xiaowei Huang

专题命中 安全评测 :safety(title,abstract);分类 cs.AI、cs.LG

Comments Preprint Accepted by ACM Transactions on Embedded Computing Systems

Journal ref ACM Transactions on Embedded Computing Systems; 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.03589 2023-01-11 eess.IV cs.AI cs.CV cs.LG 81%

Explainable, Physics Aware, Trustworthy AI Paradigm Shift for Synthetic Aperture Radar

Mihai Datcu, Zhongling Huang, Andrei Anghel, Juanping Zhao, Remus Cacoveanu

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.00951 2023-01-04 cs.CY cs.AI 81%

Digital Engineering Transformation with Trustworthy AI towards Industry 4.0: Emerging Paradigm Shifts

Jingwei Huang

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI、cs.CY

Comments Accepted Version 23 pages, 9 figures

Journal ref Transactions of the SDPS: Journal of Integrated Design and Process Science, 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.10045 2022-10-19 cs.CL cs.AI 81%

SafeText: A Benchmark for Exploring Physical Safety in Language Models

Sharon Levy, Emily Allaway, Melanie Subbiah, Lydia Chilton, Desmond Patton, Kathleen McKeown, William Yang Wang

专题命中 安全评测 :safety(title,abstract);分类 cs.CL、cs.AI

Comments Accepted to EMNLP 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.09239 2022-09-21 cs.LG cs.AI cs.CV 81%

Non-Imaging Medical Data Synthesis for Trustworthy AI: A Comprehensive Survey

Xiaodan Xing, Huanjun Wu, Lichao Wang, Iain Stenson, May Yong, Javier Del Ser, Simon Walsh, Guang Yang

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI、cs.LG

Comments 35 pages, Submitted to ACM Computing Surveys

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.14660 2022-09-01 cs.LG cs.AI cs.CV cs.RO cs.SE 81%

Unifying Evaluation of Machine Learning Safety Monitors

Joris Guerin, Raul Sena Ferreira, Kevin Delmas, Jérémie Guiochet

专题命中 安全评测 :safety(title,abstract);分类 cs.AI、cs.LG

Comments 9 pages, 5 figures, 3 tables, to appear in the proceedings of the 33rd IEEE International Symposium on Software Reliability Engineering (ISSRE 2022)

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.00898 2022-08-02 cs.LG cs.AI cs.CV 81%

Joint covariate-alignment and concept-alignment: a framework for domain generalization

Thuan Nguyen, Boyang Lyu, Prakash Ishwar, Matthias Scheutz, Shuchin Aeron

专题命中 安全评测 :alignment(title,abstract);分类 cs.AI、cs.LG

Comments 8 pages, 2 figures, and 1 table. This paper is accepted at 32nd IEEE International Workshop on Machine Learning for Signal Processing (MLSP 2022)

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.02009 2022-07-06 cs.LG cs.AI eess.SP 81%

Towards trustworthy Energy Disaggregation: A review of challenges, methods and perspectives for Non-Intrusive Load Monitoring

Maria Kaselimi, Eftychios Protopapadakis, Athanasios Voulodimos, Nikolaos Doulamis, Anastasios Doulamis

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.11981 2022-06-27 cs.AI cs.CY 81%

Never trust, always verify : a roadmap for Trustworthy AI?

Lionel Nganyewou Tidjon, Foutse Khomh

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2105.06591 2022-06-22 cs.SI cs.CY cs.LG 81%

An Interpretable Graph-based Mapping of Trustworthy Machine Learning Research

Noemi Derzsy, Subhabrata Majumdar, Rajat Malik

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.CY、cs.LG

Comments Accepted in CompleNet-2021 (oral presentation)

Journal ref In: Teixeira, A.S., Pacheco, D., Oliveira, M., Barbosa, H., Gonçalves, B., Menezes, R. (eds) Complex Networks XII. CompleNet-Live 2021. Springer Proceedings in Complexity. Springer, Cham

详情

展开后加载摘要…

URL PDF HTML 收藏
2201.13229 2022-06-14 cs.CV cs.AI cs.LG cs.SI 81%

Network-level Safety Metrics for Overall Traffic Safety Assessment: A Case Study

Xiwen Chen, Hao Wang, Abolfazl Razi, Brendan Russo, Jason Pacheco, John Roberts, Jeffrey Wishart, Larry Head, Alonso Granados Baca

专题命中 安全评测 :safety(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.01167 2022-05-27 cs.AI cs.LG 81%

Trustworthy AI: From Principles to Practices

Bo Li, Peng Qi, Bo Liu, Shuai Di, Jingen Liu, Jiquan Pei, Jinfeng Yi, Bowen Zhou

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.13373 2022-05-10 cs.HC cs.AI cs.LG 81%

Trustworthy AI and Robotics and the Implications for the AEC Industry: A Systematic Literature Review and Future Potentials

Newsha Emaminejad, Reza Akhavian

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2201.03413 2022-04-27 cs.AI cs.CY 81%

Systems Challenges for Trustworthy Embodied Systems

Harald Rueß

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI、cs.CY

Comments 57 pages, 7 figures, 3 tables. Public project deliverable, fortiss whitepaper

详情

展开后加载摘要…

URL PDF HTML 收藏
2106.06046 2022-04-13 cs.LG cs.AI cs.CR 81%

Information Theoretic Evaluation of Privacy-Leakage, Interpretability, and Transferability for Trustworthy AI

Mohit Kumar, Bernhard A. Moser, Lukas Fischer, Bernhard Freudenthaler

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI、cs.LG

Comments arXiv admin note: text overlap with arXiv:2105.04615, arXiv:2104.07060

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.06228 2022-03-15 cs.CL cs.AI 81%

CoDA21: Evaluating Language Understanding Capabilities of NLP Models With Context-Definition Alignment

Lütfi Kerem Senel, Timo Schick, Hinrich Schütze

专题命中 安全评测 :alignment(title,abstract);分类 cs.CL、cs.AI

Comments To appear in ACL 2022, 5 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2202.07447 2022-03-11 cs.CY cs.AI 81%

Trustworthy Autonomous Systems (TAS): Engaging TAS experts in curriculum design

Mohammad Naiseh, Caitlin Bentley, Sarvapali D. Ramchurn

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.03718 2022-03-09 cs.CY cs.AI 81%

Towards User-Centered Metrics for Trustworthy AI in Immersive Cyberspace

Pengyuan Zhou, Benjamin Finley, Lik-Hang Lee, Yong Liao, Haiyong Xie, Pan Hui

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2009.06704 2022-02-16 cs.LG cs.CL 81%

Food safety risk prediction with Deep Learning models using categorical embeddings on European Union data

Alberto Nogales, Rodrigo Díaz Morón, Álvaro J. García-Tejedor

专题命中 安全评测 :safety(title,abstract);分类 cs.CL、cs.LG

Comments 20 pages,8 figures, 15 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.07773 2021-12-16 cs.AI cs.CY 81%

Filling gaps in trustworthy development of AI

Shahar Avin, Haydn Belfield, Miles Brundage, Gretchen Krueger, Jasmine Wang, Adrian Weller, Markus Anderljung, Igor Krawczuk, David Krueger, Jonathan Lebensold, Tegan Maharaj, Noa Zilberman

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI、cs.CY

Journal ref Science (2021) Vol 374, Issue 6573, pp. 1327-1329

详情

展开后加载摘要…

URL PDF HTML 收藏
2108.06912 2021-11-01 cs.LG cs.AI 81%

Blockchain-based Trustworthy Federated Learning Architecture

Sin Kit Lo, Yue Liu, Qinghua Lu, Chen Wang, Xiwei Xu, Hye-Young Paik, Liming Zhu

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2108.11844 2021-08-27 cs.CY cs.AI 81%

AI at work -- Mitigating safety and discriminatory risk with technical standards

Nikolas Becker, Pauline Junginger, Lukas Martinez, Daniel Krupka, Leonie Beining

专题命中 安全评测 :safety(title,abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2107.07334 2021-07-16 cs.HC cs.CR cs.CY cs.LG 81%

Tournesol: A quest for a large, secure and trustworthy database of reliable human judgments

Lê-Nguyên Hoang, Louis Faucon, Aidan Jungo, Sergei Volodin, Dalia Papuc, Orfeas Liossatos, Ben Crulis, Mariame Tighanimine, Isabela Constantin, Anastasiia Kucherenko, Alexandre Maurer, Felix Grimberg, Vlad Nitu, Chris Vossen, Sébastien Rouault, El-Mahdi El-Mhamdi

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.CY、cs.LG

Comments 27 pages, 13 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2103.09051 2021-03-17 cs.CY cs.AI cs.SE 81%

Exploring the Assessment List for Trustworthy AI in the Context of Advanced Driver-Assistance Systems

Markus Borg, Joshua Bronson, Linus Christensson, Fredrik Olsson, Olof Lennartsson, Elias Sonnsjö, Hamid Ebabi, Martin Karsberg

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI、cs.CY

Comments Accepted for publication in the Proc. of the 2nd Workshop on Ethics in Software Engineering Research and Practice

详情

展开后加载摘要…

URL PDF HTML 收藏
2101.05620 2021-01-15 cs.LG cs.CY 81%

A Framework for Assurance of Medication Safety using Machine Learning

Yan Jia, Tom Lawton, John McDermid, Eric Rojas, Ibrahim Habli

专题命中 安全评测 :safety(title,abstract);分类 cs.CY、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2007.15911 2021-01-06 cs.AI cs.LG stat.ML 81%

The role of explainability in creating trustworthy artificial intelligence for health care: a comprehensive survey of the terminology, design choices, and evaluation strategies

Aniek F. Markus, Jan A. Kors, Peter R. Rijnbeek

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI、cs.LG

Journal ref Journal of Biomedical Informatics, 113 (2021), 103655

详情

展开后加载摘要…

URL PDF HTML 收藏
2012.06373 2020-12-14 cs.LG cs.AI cs.AR cs.NE stat.ML 81%

Hardware Beyond Backpropagation: a Photonic Co-Processor for Direct Feedback Alignment

Julien Launay, Iacopo Poli, Kilian Müller, Gustave Pariente, Igor Carron, Laurent Daudet, Florent Krzakala, Sylvain Gigan

专题命中 安全评测 :alignment(title,abstract);分类 cs.AI、cs.LG

Comments 6 pages, 2 figures, 1 table. Oral at the Beyond Backpropagation Workshop, NeurIPS 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2011.02272 2020-11-05 cs.CY cs.CR cs.CV cs.LG 81%

Trustworthy AI

Richa Singh, Mayank Vatsa, Nalini Ratha

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.CY、cs.LG

Comments ACM CODS-COMAD 2021 Tutorial

详情

展开后加载摘要…

URL PDF HTML 收藏
2007.07768 2020-09-01 cs.SE cs.AI cs.CY 81%

Opening the Software Engineering Toolbox for the Assessment of Trustworthy AI

Mohit Kumar Ahuja, Mohamed-Bachir Belaid, Pierre Bernabé, Mathieu Collet, Arnaud Gotlieb, Chhagan Lal, Dusica Marijan, Sagar Sen, Aizaz Sharif, Helge Spieker

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI、cs.CY

Comments 1st International Workshop on New Foundations for Human-Centered AI @ ECAI 2020

详情

展开后加载摘要…

URL PDF HTML 收藏