arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 9380 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 9380 篇

2305.02368 2023-05-05 cs.LG cs.AI 62%

Metric Tools for Sensitivity Analysis with Applications to Neural Networks

Jaime Pizarroso, David Alfaya, José Portela, Antonio Muñoz

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

Comments 15 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.13714 2023-05-02 cs.AI cs.CL cs.IR 62%

Evaluation of GPT-3.5 and GPT-4 for supporting real-world information needs in healthcare delivery

Debadutta Dash, Rahul Thapa, Juan M. Banda, Akshay Swaminathan, Morgan Cheatham, Mehr Kashyap, Nikesh Kotecha, Jonathan H. Chen, Saurabh Gombar, Lance Downing, Rachel Pedreira, Ethan Goh, Angel Arnaout, Garret Kenn Morris, Honor Magon, Matthew P Lungren, Eric Horvitz, Nigam H. Shah

专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.AI

Comments 27 pages including supplemental information

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.00278 2023-05-02 cs.CV cs.AI cs.LG 62%

Segment Anything Model (SAM) Meets Glass: Mirror and Transparent Objects Cannot Be Easily Detected

Dongsheng Han, Chaoning Zhang, Yu Qiao, Maryam Qamar, Yuna Jung, SeungKyu Lee, Sung-Ho Bae, Choong Seon Hong

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.09056 2023-04-27 cs.LG cs.AI 62%

Concept Embedding Models: Beyond the Accuracy-Explainability Trade-Off

Mateo Espinosa Zarlenga, Pietro Barbiero, Gabriele Ciravegna, Giuseppe Marra, Francesco Giannini, Michelangelo Diligenti, Zohreh Shams, Frederic Precioso, Stefano Melacci, Adrian Weller, Pietro Lio, Mateja Jamnik

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

Comments To appear at NeurIPS 2022

Journal ref https://proceedings.neurips.cc/paper_files/paper/2022/hash/867c06823281e506e8059f5c13a57f75-Abstract-Conference.html

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.05524 2023-04-13 cs.LG cs.CL 62%

Understanding Causality with Large Language Models: Feasibility and Opportunities

Cheng Zhang, Stefan Bauer, Paul Bennett, Jiangfeng Gao, Wenbo Gong, Agrin Hilmkil, Joel Jennings, Chao Ma, Tom Minka, Nick Pawlowski, James Vaughan

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.13375 2023-04-13 cs.CL cs.AI 62%

Capabilities of GPT-4 on Medical Challenge Problems

Harsha Nori, Nicholas King, Scott Mayer McKinney, Dean Carignan, Eric Horvitz

专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.AI

Comments 35 pages, 15 figures; added GPT-4-base model results and discussion

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.04780 2023-04-12 cs.LG cs.AI 62%

A Review on Explainable Artificial Intelligence for Healthcare: Why, How, and When?

Subrato Bharati, M. Rubaiyat Hossain Mondal, Prajoy Podder

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

Comments 15 pages, 3 figures, accepted for publication in the IEEE Transactions on Artificial Intelligence

Journal ref IEEE Transactions on Artificial Intelligence, 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.13567 2023-03-10 cs.LG cs.AI cs.CR cs.CV 62%

Towards Audit Requirements for AI-based Systems in Mobility Applications

Devi Padmavathi Alagarswamy, Christian Berghoff, Vasilios Danos, Fabian Langer, Thora Markert, Georg Schneider, Arndt von Twickel, Fabian Woitschek

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG

Comments To appear in Proceedings of the 9th International Conference on Information Systems Security and Privacy

Journal ref Proceedings of the 9th International Conference on Information Systems Security and Privacy - ICISSP, pp. 339-348, 2023 , Lisbon, Portugal

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.08955 2023-03-08 cs.CY cs.HC cs.LG 62%

Trusting the Explainers: Teacher Validation of Explainable Artificial Intelligence for Course Design

Vinitra Swamy, Sijia Du, Mirko Marras, Tanja Käser

专题命中 安全评测 :trustworthy(abstract);分类 cs.CY、cs.LG

Comments Accepted as a full paper (Best Paper nominee) at LAK 2023: The 13th International Learning Analytics and Knowledge Conference, March 13-17, 2023, Arlington, Texas, USA

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.03185 2023-03-07 cs.CV cs.AI cs.LG 62%

Evaluation of Confidence-based Ensembling in Deep Learning Image Classification

Rafael Rosales, Peter Popov, Michael Paulitsch

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.13122 2023-03-06 cs.CR cs.AI cs.LG 62%

Towards Adversarial Realism and Robust Learning for IoT Intrusion Detection and Classification

João Vitorino, Isabel Praça, Eva Maia

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

Comments 19 pages, 5 tables, 7 figures, Annals of Telecommunications journal

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.02080 2023-02-27 cs.AI cs.HC cs.LG 62%

Semantic match: Debugging feature attribution methods in XAI for healthcare

Giovanni Cinà, Tabea E. Röber, Rob Goedhart, Ş. İlker Birbil

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.04105 2023-02-24 cs.CL cs.LG stat.ML 62%

Words are all you need? Language as an approximation for human similarity judgments

Raja Marjieh, Pol van Rijn, Ilia Sucholutsky, Theodore R. Sumers, Harin Lee, Thomas L. Griffiths, Nori Jacoby

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.LG

Comments Accepted to ICLR 2023, final revision. https://openreview.net/forum?id=O-G91-4cMdv

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.10671 2023-02-22 cs.HC cs.AI cs.LG cs.SE 62%

Directive Explanations for Monitoring the Risk of Diabetes Onset: Introducing Directive Data-Centric Explanations and Combinations to Support What-If Explorations

Aditya Bhattacharya, Jeroen Ooge, Gregor Stiglic, Katrien Verbert

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

Comments \c{opyright} Bhattacharya et al, 2023. This is the author's version of the work. It is posted here for your personal use. Not for redistribution. Copyright is held by the owner/author(s). Publication rights licensed to ACM. The definitive version was published in ACM IUI '23: 28th International Conference on Intelligent User Interfaces Proceedings, https://doi.org/10.1145/3581641.3584075

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.10281 2023-02-22 cs.CV cs.AI cs.CL 62%

LiT Tuned Models for Efficient Species Detection

Andre Nakkab, Benjamin Feuer, Chinmay Hegde

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

Comments 5 pages, 5 figures, 1 table, presented at AAAI 2023 conference for the AIAFS workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.04844 2023-02-10 cs.CY cs.AI 62%

The Gradient of Generative AI Release: Methods and Considerations

Irene Solaiman

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.03037 2023-02-08 cs.HC cs.AI cs.LG 62%

LiteVR: Interpretable and Lightweight Cybersickness Detection using Explainable AI

Ripan Kumar Kundu, Rifatul Islam, John Quarles, Khaza Anuarul Hoque

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

Comments Accepted for publication in IEEE VR 2023 conference. arXiv admin note: substantial text overlap with arXiv:2302.01985

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.01985 2023-02-07 cs.LG cs.AI cs.HC 62%

VR-LENS: Super Learning-based Cybersickness Detection and Explainable AI-Guided Deployment in Virtual Reality

Ripan Kumar Kundu, Osama Yahia Elsaid, Prasad Calyam, Khaza Anuarul Hoque

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

Comments Accepted for publication in IUI 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.12151 2023-01-31 cs.LG cs.AI 62%

Selecting Models based on the Risk of Damage Caused by Adversarial Attacks

Jona Klemenc, Holger Trittenbach

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.00541 2023-01-31 cs.AI cs.LG 62%

Logic-Based Explainability in Machine Learning

Joao Marques-Silva

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.04046 2023-01-30 cs.CV cs.AI cs.LG 62%

Sparse Mixture-of-Experts are Domain Generalizable Learners

Bo Li, Yifei Shen, Jingkang Yang, Yezhen Wang, Jiawei Ren, Tong Che, Jun Zhang, Ziwei Liu

专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.LG

Comments ICLR 2023 (accepted as Oral presentation)

详情

展开后加载摘要…

URL PDF HTML 收藏
2202.07201 2023-01-06 cs.LG cs.AI cs.CR 62%

Holistic Adversarial Robustness of Deep Learning Models

Pin-Yu Chen, Sijia Liu

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG

Comments survey paper on holistic adversarial robustness for deep learning; published at AAAI 2023 Senior Member Presentation Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2105.10197 2022-12-23 cs.LG cs.AI 62%

Yes We Care! -- Certification for Machine Learning Methods through the Care Label Framework

Katharina Morik, Helena Kotthaus, Raphael Fischer, Sascha Mücke, Matthias Jakobs, Nico Piatkowski, Andreas Pauly, Lukas Heppe, Danny Heinrich

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

Journal ref Frontiers in Artificial Intelligence, September 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.05030 2022-12-22 cs.AI cs.HC cs.LG 62%

Assessing the communication gap between AI models and healthcare professionals: explainability, utility and trust in AI-driven clinical decision-making

Oskar Wysocki, Jessica Katharine Davies, Markel Vigo, Anne Caroline Armstrong, Dónal Landers, Rebecca Lee, André Freitas

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG

Comments supplementary information in the main pdf

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.01738 2022-11-29 quant-ph cond-mat.dis-nn cs.AI cs.LG 62%

Experimental quantum adversarial learning with programmable superconducting qubits

Wenhui Ren, Weikang Li, Shibo Xu, Ke Wang, Wenjie Jiang, Feitong Jin, Xuhao Zhu, Jiachen Chen, Zixuan Song, Pengfei Zhang, Hang Dong, Xu Zhang, Jinfeng Deng, Yu Gao, Chuanyu Zhang, Yaozu Wu, Bing Zhang, Qiujiang Guo, Hekang Li, Zhen Wang, Jacob Biamonte, Chao Song, Dong-Ling Deng, H. Wang

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG

Comments 26 pages, 17 figures, 8 algorithms

Journal ref Nature Computational Science 2, 711 (2022)

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.00639 2022-11-29 cs.CV cs.AI cs.LG 62%

A Systematic Review of Robustness in Deep Learning for Computer Vision: Mind the gap?

Nathan Drenkow, Numair Sani, Ilya Shpitser, Mathias Unberath

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2010.10596 2022-11-17 cs.LG cs.AI stat.ML 62%

Counterfactual Explanations and Algorithmic Recourses for Machine Learning: A Review

Sahil Verma, Varich Boonsanong, Minh Hoang, Keegan E. Hines, John P. Dickerson, Chirag Shah

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

Comments 23 pages (8 pages of references)

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.07941 2022-11-16 cs.RO cs.AI cs.LG 62%

Automatic Evaluation of Excavator Operators using Learned Reward Functions

Pranav Agarwal, Marek Teichmann, Sheldon Andrews, Samira Ebrahimi Kahou

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG

Comments 11 pages, 5 figures, Accepted at Reinforcement Learning for Real Life (RL4RealLife) Workshop at NeurIPS 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.01036 2022-11-08 cs.AI cs.HC cs.LG 62%

Explainable AI over the Internet of Things (IoT): Overview, State-of-the-Art and Future Directions

Senthil Kumar Jagatheesaperumal, Quoc-Viet Pham, Rukhsana Ruby, Zhaohui Yang, Chunmei Xu, Zhaoyang Zhang

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

Comments 29 pages, 7 figures, 2 tables. IEEE Open Journal of the Communications Society (2022)

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.00004 2022-10-28 cs.CL cs.LG 62%

Reproducibility Issues for BERT-based Evaluation Metrics

Yanran Chen, Jonas Belouadi, Steffen Eger

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.LG

Comments EMNLP 2022 Camera-Ready (captions fixed)

详情

展开后加载摘要…

URL PDF HTML 收藏