arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 9380 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 9380 篇

2209.14613 2023-09-04 cs.LG cs.CY 62%

Fair admission risk prediction with proportional multicalibration

William La Cava, Elle Lett, Guangya Wan

专题命中 安全评测 :trustworthy(abstract);分类 cs.CY、cs.LG

Comments Published in the 2023 Conference on Health, Inference, and Learning (CHIL). Best paper award

Journal ref Proceedings of Machine Learning Research 209 (2023) 350-378

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.17926 2023-08-31 cs.CL cs.AI cs.IR 62%

Large Language Models are not Fair Evaluators

Peiyi Wang, Lei Li, Liang Chen, Zefan Cai, Dawei Zhu, Binghuai Lin, Yunbo Cao, Qi Liu, Tianyu Liu, Zhifang Sui

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.11638 2023-08-24 eess.SP cs.AI cs.LG 62%

IoT Data Trust Evaluation via Machine Learning

Timothy Tadj, Reza Arablouei, Volkan Dedeoglu

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.12031 2023-08-21 cs.CL cs.AI 62%

Clinical Camel: An Open Expert-Level Medical Language Model with Dialogue-Based Knowledge Encoding

Augustin Toma, Patrick R. Lawler, Jimmy Ba, Rahul G. Krishnan, Barry B. Rubin, Bo Wang

专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.AI

Comments for model weights, see https://huggingface.co/wanglab/

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.07902 2023-08-16 cs.CL cs.AI 62%

Through the Lens of Core Competency: Survey on Evaluation of Large Language Models

Ziyu Zhuang, Qiguang Chen, Longxuan Ma, Mingda Li, Yi Han, Yushan Qian, Haopeng Bai, Zixian Feng, Weinan Zhang, Ting Liu

专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.14961 2023-08-11 cs.LG cs.AI cs.CV 62%

Diffusion Denoised Smoothing for Certified and Adversarial Robust Out-Of-Distribution Detection

Nicola Franco, Daniel Korth, Jeanette Miriam Lorenz, Karsten Roscher, Stephan Guennemann

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.01840 2023-08-04 cs.LG cs.AI cs.CR 62%

URET: Universal Robustness Evaluation Toolkit (for Evasion)

Kevin Eykholt, Taesung Lee, Douglas Schales, Jiyong Jang, Ian Molloy, Masha Zorin

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG

Comments Accepted at USENIX '23

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.09418 2023-08-01 cs.LG cs.AI 62%

SAFARI: Versatile and Efficient Evaluations for Robustness of Interpretability

Wei Huang, Xingyu Zhao, Gaojie Jin, Xiaowei Huang

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

Comments Accepted by the IEEE/CVF International Conference on Computer Vision 2023 (ICCV'23)

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.12518 2023-07-25 cs.LG cs.AI cs.IR 62%

FaFCNN: A General Disease Classification Framework Based on Feature Fusion Neural Networks

Menglin Kong, Shaojie Zhao, Juan Cheng, Xingquan Li, Ri Su, Muzhou Hou, Cong Cao

专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.03447 2023-07-25 cs.AI cs.LG q-bio.GN 62%

Machine Learning-Friendly Biomedical Datasets for Equivalence and Subsumption Ontology Matching

Yuan He, Jiaoyan Chen, Hang Dong, Ernesto Jiménez-Ruiz, Ali Hadian, Ian Horrocks

专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.LG

Comments Accepted paper (Best Resource Paper Candidate) in the 21st International Semantic Web Conference (ISWC-2022); Bio-ML Dataset: https://doi.org/10.5281/zenodo.6510086

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.14102 2023-07-10 cs.RO cs.AI cs.LG cs.MA 62%

SocNavGym: A Reinforcement Learning Gym for Social Navigation

Aditya Kapoor, Sushant Swamy, Luis Manso, Pilar Bachiller

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG

Comments IEEE RO-MAN

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.02971 2023-07-07 cs.CV cs.AI cs.CL 62%

On the Cultural Gap in Text-to-Image Generation

Bingshuai Liu, Longyue Wang, Chenyang Lyu, Yong Zhang, Jinsong Su, Shuming Shi, Zhaopeng Tu

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

Comments Equal contribution: Bingshuai Liu and Longyue Wang. Work done while Bingshuai Liu and Chengyang Lyu were interning at Tencent AI Lab. Zhaopeng Tu is the corresponding author

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.02502 2023-07-07 q-bio.OT cs.AI cs.CL 62%

Math Agents: Computational Infrastructure, Mathematical Embedding, and Genomics

Melanie Swan, Takashi Kido, Eric Roland, Renato P. dos Santos

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.13671 2023-06-27 cs.CY cs.AI cs.HC 62%

Deceptive AI Ecosystems: The Case of ChatGPT

Xiao Zhan, Yifan Xu, Stefan Sarkadi

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.CY

Comments 6 pages, To appear in the Proceedings of the 2023 ACM conference on Conversational User Interfaces (CUI 23)

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.11915 2023-06-27 cs.LG cs.AI cs.SI stat.ML 62%

Structure-Aware Robustness Certificates for Graph Classification

Pierre Osselin, Henry Kenlay, Xiaowen Dong

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG

Comments 9 pages, 6 figures (15 pages, 10 figures including references and appendices)

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.11985 2023-06-22 cs.LG cs.CY 62%

Evaluation of Popular XAI Applied to Clinical Prediction Models: Can They be Trusted?

Aida Brankovic, David Cook, Jessica Rahman, Wenjie Huang, Sankalp Khanna

专题命中 安全评测 :trustworthy(abstract);分类 cs.CY、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.00024 2023-06-21 cs.CL cs.LG 62%

Self-Verification Improves Few-Shot Clinical Information Extraction

Zelalem Gero, Chandan Singh, Hao Cheng, Tristan Naumann, Michel Galley, Jianfeng Gao, Hoifung Poon

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.LG

Journal ref IMLH 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.11911 2023-06-21 cs.LG cs.AI stat.ML 62%

Multi-dimensional concept discovery (MCD): A unifying framework with completeness guarantees

Johanna Vielhaben, Stefan Blücher, Nils Strodthoff

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

Comments v2: Version published by Transactions on Machine Learning Research in 2023 (TMLR ISSN 2835-8856) https://openreview.net/forum?id=KxBQPz7HKh. 25 pages, 11 figures. This work builds on an earlier manuscript (arXiv:2203.06043) and crucially extends it. Code is available at https://github.com/jvielhaben/MCD-XAI

Journal ref Version published by Transactions on Machine Learning Research in 2023 (TMLR ISSN 2835-8856) https://openreview.net/forum?id=KxBQPz7HKh

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.04757 2023-06-16 cs.CL cs.AI 62%

INSTRUCTEVAL: Towards Holistic Evaluation of Instruction-Tuned Large Language Models

Yew Ken Chia, Pengfei Hong, Lidong Bing, Soujanya Poria

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

Comments Github: https://github.com/declare-lab/instruct-eval Leaderboard: https://declare-lab.github.io/instruct-eval/

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.04454 2023-06-08 cs.LG cs.AI 62%

Training-Free Neural Active Learning with Initialization-Robustness Guarantees

Apivich Hemachandra, Zhongxiang Dai, Jasraj Singh, See-Kiong Ng, Bryan Kian Hsiang Low

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG

Comments Accepted to 40th International Conference on Machine Learning (ICML 2023), 41 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.03902 2023-06-07 cs.CL cs.AI cs.LO q-bio.NC 62%

Utterance Classification with Logical Neural Network: Explainable AI for Mental Disorder Diagnosis

Yeldar Toleubay, Don Joven Agravante, Daiki Kimura, Baihan Lin, Djallel Bouneffouf, Michiaki Tatsubori

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.AI

Comments ACL 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.03315 2023-06-07 cs.CL cs.AI 62%

Few Shot Rationale Generation using Self-Training with Dual Teachers

Aditya Srikanth Veerubhotla, Lahari Poddar, Jun Yin, György Szarvas, Sharanya Eswaran

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.AI

Comments ACL Findings 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.10965 2023-06-06 cs.CV cs.AI cs.LG 62%

CLIP-Dissect: Automatic Description of Neuron Representations in Deep Vision Networks

Tuomas Oikarinen, Tsui-Wei Weng

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

Comments Published in ICLR 2023 Conference (Spotlight). New v5(5 June 2023) - Added crowdsourced user study in Appendix B, not included in ICLR publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.03063 2023-06-06 cs.LG cs.AI 62%

Neuro-symbolic model for cantilever beams damage detection

Darian Onchis, Gilbert-Rainer Gillich, Eduard Hogea, Cristian Tufisi

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.16328 2023-05-29 cs.CL cs.LG 62%

Semantic Composition in Visually Grounded Language Models

Rohan Pandey

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.LG

Comments Carnegie Mellon University Senior Thesis. arXiv admin note: substantial text overlap with arXiv:2212.10549

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.14099 2023-05-25 cs.LG cs.AI 62%

An Explainable-AI approach for Diagnosis of COVID-19 using MALDI-ToF Mass Spectrometry

Venkata Devesh Reddy Seethi, Zane LaCasse, Prajkta Chivte, Joshua Bland, Shrihari S. Kadkol, Elizabeth R. Gaillard, Pratool Bharti, Hamed Alhoori

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.13246 2023-05-23 cs.CL cs.AI 62%

Interactive Natural Language Processing

Zekun Wang, Ge Zhang, Kexin Yang, Ning Shi, Wangchunshu Zhou, Shaochun Hao, Guangzheng Xiong, Yizhi Li, Mong Yuan Sim, Xiuying Chen, Qingqing Zhu, Zhenzhu Yang, Adam Nik, Qi Liu, Chenghua Lin, Shi Wang, Ruibo Liu, Wenhu Chen, Ke Xu, Dayiheng Liu, Yike Guo, Jie Fu

专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.AI

Comments 110 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.10784 2023-05-19 cs.CL cs.HC cs.LG 62%

Eyettention: An Attention-based Dual-Sequence Model for Predicting Human Scanpaths during Reading

Shuwen Deng, David R. Reich, Paul Prasse, Patrick Haller, Tobias Scheffer, Lena A. Jäger

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.11879 2023-05-19 cs.AI cs.CL 62%

Case-Based Reasoning with Language Models for Classification of Logical Fallacies

Zhivar Sourati, Filip Ilievski, Hông-Ân Sandlin, Alain Mermoud

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.05750 2023-05-11 cs.LG cs.AI cs.AR 62%

A Systematic Literature Review on Hardware Reliability Assessment Methods for Deep Neural Networks

Mohammad Hasan Ahmadilivani, Mahdi Taheri, Jaan Raik, Masoud Daneshtalab, Maksim Jenihhin

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG

Comments 42 pages, 15 figures, 3 tables, 201 references. The paper has been reviewed and revised 2 times and is under the 3rd review

详情

展开后加载摘要…

URL PDF HTML 收藏