arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 9434 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 9434 篇

2308.04030 2023-08-09 cs.AI 57%

Gentopia: A Collaborative Platform for Tool-Augmented LLMs

Binfeng Xu, Xukun Liu, Hua Shen, Zeyu Han, Yuhan Li, Murong Yue, Zhiyuan Peng, Yuchen Liu, Ziyu Yao, Dongkuan Xu

专题命中 安全评测 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.00500 2023-08-09 cs.CV cs.LG eess.IV 57%

Inherently Interpretable Multi-Label Classification Using Class-Specific Counterfactuals

Susu Sun, Stefano Woerner, Andreas Maier, Lisa M. Koch, Christian F. Baumgartner

专题命中 安全评测 :trustworthy(abstract);分类 cs.LG

Comments Accepted to MIDL 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.03638 2023-08-08 cs.CL 57%

KITLM: Domain-Specific Knowledge InTegration into Language Models for Question Answering

Ankush Agarwal, Sakharam Gawade, Amar Prakash Azad, Pushpak Bhattacharyya

专题命中 安全评测 :safety(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.03179 2023-08-08 cs.AI 57%

Empirical Optimal Risk to Quantify Model Trustworthiness for Failure Detection

Shuang Ao, Stefan Rueger, Advaith Siddharthan

专题命中 安全评测 :safety(abstract);分类 cs.AI

Comments 7 pages

Journal ref 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2201.12501 2023-08-01 cs.CL 57%

Does Transliteration Help Multilingual Language Modeling?

Ibraheem Muhammad Moosa, Mahmud Elahi Akhter, Ashfia Binte Habib

专题命中 安全评测 :alignment(abstract);分类 cs.CL

Comments In Findings of the Association for Computational Linguistics: EACL 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.15452 2023-07-31 cs.HC cs.CY 57%

From OECD to India: Exploring cross-cultural differences in perceived trust, responsibility and reliance of AI and human experts

Vishakha Agrawal, Serhiy Kandul, Markus Kneer, Markus Christen

专题命中 安全评测 :trustworthy(abstract);分类 cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.11558 2023-07-24 cs.CV cs.CL 57%

Advancing Visual Grounding with Scene Knowledge: Benchmark and Method

Zhihong Chen, Ruifei Zhang, Yibing Song, Xiang Wan, Guanbin Li

专题命中 安全评测 :alignment(abstract);分类 cs.CL

Comments Computer Vision and Natural Language Processing. 21 pages, 14 figures. CVPR-2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.10714 2023-07-21 eess.SY cs.AI cs.RO cs.SY 57%

Introducing Risk Shadowing For Decisive and Comfortable Behavior Planning

Tim Puphal, Julian Eggert

专题命中 安全评测 :safety(abstract);分类 cs.AI

Comments Accepted at IEEE ITSC 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.03719 2023-07-19 cs.LG 57%

A survey on learning from imbalanced data streams: taxonomy, challenges, empirical study, and reproducible experimental framework

Gabriel Aguiar, Bartosz Krawczyk, Alberto Cano

专题命中 安全评测 :trustworthy(abstract);分类 cs.LG

Journal ref Machine Learning, 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.02412 2023-07-06 cs.CR cs.AI 57%

Android Malware Detection using Machine learning: A Review

Md Naseef-Ur-Rahman Chowdhury, Ahshanul Haque, Hamdy Soliman, Mohammad Sahinur Hossen, Tanjim Fatima, Imtiaz Ahmed

专题命中 安全评测 :safety(abstract);分类 cs.AI

Comments 22 pages,2 figures, IntelliSys 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.17323 2023-07-04 cs.LG 57%

Scaling Model Checking for DNN Analysis via State-Space Reduction and Input Segmentation (Extended Version)

Mahum Naseer, Osman Hasan, Muhammad Shafique

专题命中 安全评测 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2001.02658 2023-07-04 cs.LG cs.CV 57%

Distributionally Robust Deep Learning using Hardness Weighted Sampling

Lucas Fidon, Michael Aertsen, Thomas Deprest, Doaa Emam, Frédéric Guffens, Nada Mufti, Esther Van Elslander, Ernst Schwartz, Michael Ebner, Daniela Prayer, Gregor Kasprian, Anna L. David, Andrew Melbourne, Sébastien Ourselin, Jan Deprest, Georg Langs, Tom Vercauteren

专题命中 安全评测 :safety(abstract);分类 cs.LG

Comments Accepted for publication at the Journal of Machine Learning for Biomedical Imaging (MELBA) https://www.melba-journal.org/papers/2022:019.html

Journal ref https://www.melba-journal.org/papers/2022:019.html

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.16774 2023-06-30 cs.CL 57%

Stop Pre-Training: Adapt Visual-Language Models to Unseen Languages

Yasmine Karoui, Rémi Lebret, Negar Foroutan, Karl Aberer

专题命中 安全评测 :alignment(abstract);分类 cs.CL

Comments Accepted to ACL 2023 as short paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.15159 2023-06-28 stat.ML cs.LG math.DS 57%

Evaluation of machine learning architectures on the quantification of epistemic and aleatoric uncertainties in complex dynamical systems

Stephen Guth, Alireza Mojahed, Themistoklis P. Sapsis

专题命中 安全评测 :safety(abstract);分类 cs.LG

Comments Submitted for publication to "Computer Methods in Applied Mechanics and Engineering." 25 pages, 20 figures. arXiv admin note: text overlap with arXiv:1505.05424 by other authors

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.11543 2023-06-28 cs.CL 57%

Constructing Word-Context-Coupled Space Aligned with Associative Knowledge Relations for Interpretable Language Modeling

Fanyu Wang, Zhenping Xie

专题命中 安全评测 :alignment(abstract);分类 cs.CL

Comments Accepted at ACL 2023, Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.15225 2023-06-27 cs.CL 57%

SAIL: Search-Augmented Instruction Learning

Hongyin Luo, Yung-Sung Chuang, Yuan Gong, Tianhua Zhang, Yoon Kim, Xixin Wu, Danny Fox, Helen Meng, James Glass

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.13501 2023-06-26 cs.CL 57%

Knowledge-Infused Self Attention Transformers

Kaushik Roy, Yuxin Zi, Vignesh Narayanan, Manas Gaur, Amit Sheth

专题命中 安全评测 :alignment(abstract);分类 cs.CL

Comments Accepted for publication at the Second Workshop on Knowledge Augmented Methods for NLP, colocated with KDD 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.09989 2023-06-22 cs.LG stat.ML 57%

Finding Competence Regions in Domain Generalization

Jens Müller, Stefan T. Radev, Robert Schmier, Felix Draxler, Carsten Rother, Ullrich Köthe

专题命中 安全评测 :trustworthy(abstract);分类 cs.LG

Comments The paper has been published at TMLR (see https://openreview.net/forum?id=TSy0vuwQFN)

Journal ref Transactions on Machine Learning Research (06/2023)

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.08682 2023-06-16 stat.ML cs.LG 57%

Predicting Real-time Crash Risks during Hurricane Evacuation Using Connected Vehicle Data

Zaheen E Muktadi Syed, Samiul Hasan

专题命中 安全评测 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.08238 2023-06-16 cs.HC cs.AI 57%

Maestro: A Gamified Platform for Teaching AI Robustness

Margarita Geleta, Jiacen Xu, Manikanta Loya, Junlin Wang, Sameer Singh, Zhou Li, Sergio Gago-Masague

专题命中 安全评测 :safety(abstract);分类 cs.AI

Comments 9 pages, 6 figures, published at the Thirteenth Symposium on Educational Advances in Artificial Intelligence (EAAI-23) in the Association for the Advancement of Artificial Intelligence Conference (AAAI), 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.04717 2023-06-13 cs.CV cs.AI eess.IV 57%

AGIQA-3K: An Open Database for AI-Generated Image Quality Assessment

Chunyi Li, Zicheng Zhang, Haoning Wu, Wei Sun, Xiongkuo Min, Xiaohong Liu, Guangtao Zhai, Weisi Lin

专题命中 安全评测 :alignment(abstract);分类 cs.AI

Comments 12 pages, 11 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2104.01393 2023-06-13 cs.CL eess.AS 57%

On-the-Fly Aligned Data Augmentation for Sequence-to-Sequence ASR

Tsz Kin Lam, Mayumi Ohta, Shigehiko Schamoni, Stefan Riezler

专题命中 安全评测 :alignment(abstract);分类 cs.CL

Comments Accepted at INTERSPEECH 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.14174 2023-06-13 math.OC cs.LG 57%

Explainable AI via Learning to Optimize

Howard Heaton, Samy Wu Fung

专题命中 安全评测 :trustworthy(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.04362 2023-06-08 cs.CV cs.CL 57%

Youku-mPLUG: A 10 Million Large-scale Chinese Video-Language Dataset for Pre-training and Benchmarks

Haiyang Xu, Qinghao Ye, Xuan Wu, Ming Yan, Yuan Miao, Jiabo Ye, Guohai Xu, Anwen Hu, Yaya Shi, Guangwei Xu, Chenliang Li, Qi Qian, Maofei Que, Ji Zhang, Xiao Zeng, Fei Huang

专题命中 安全评测 :safety(abstract);分类 cs.CL

Comments Working in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.03980 2023-06-08 cs.AI 57%

Counterfactual Explanations and Predictive Models to Enhance Clinical Decision-Making in Schizophrenia using Digital Phenotyping

Juan Sebastian Canas, Francisco Gomez, Omar Costilla-Reyes

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.03802 2023-06-07 cs.CV cs.AI 57%

Learning to Ground Instructional Articles in Videos through Narrations

Effrosyni Mavroudi, Triantafyllos Afouras, Lorenzo Torresani

专题命中 安全评测 :alignment(abstract);分类 cs.AI

Comments 17 pages, 4 figures and 10 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.10838 2023-06-06 cs.LG cs.PL 57%

ProgSG: Cross-Modality Representation Learning for Programs in Electronic Design Automation

Yunsheng Bai, Atefeh Sohrabizadeh, Zongyue Qin, Ziniu Hu, Yizhou Sun, Jason Cong

专题命中 安全评测 :alignment(abstract);分类 cs.LG

Comments Requires further polishing

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.05011 2023-06-02 cs.HC cs.CL 57%

Towards an Understanding and Explanation for Mixed-Initiative Artificial Scientific Text Detection

Luoxuan Weng, Minfeng Zhu, Kam Kwai Wong, Shi Liu, Jiashun Sun, Hang Zhu, Dongming Han, Wei Chen

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.19636 2023-06-01 cs.LG 57%

Explainable AI for Malnutrition Risk Prediction from m-Health and Clinical Data

Flavio Di Martino, Franca Delmastro, Cristina Dolciotti

专题命中 安全评测 :trustworthy(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.17499 2023-05-30 cs.CL cs.MM eess.AS 57%

CIF-PT: Bridging Speech and Text Representations for Spoken Language Understanding via Continuous Integrate-and-Fire Pre-Training

Linhao Dong, Zhecheng An, Peihao Wu, Jun Zhang, Lu Lu, Zejun Ma

专题命中 安全评测 :alignment(abstract);分类 cs.CL

Comments Accepted by ACL 2023 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏