arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 9434 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 9434 篇

2204.07532 2022-09-20 q-bio.BM cs.LG q-bio.QM 57%

Accurate ADMET Prediction with XGBoost

Hao Tian, Rajas Ketkar, Peng Tao

专题命中 安全评测 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.07572 2022-09-19 cs.CY 57%

Power to the People? Opportunities and Challenges for Participatory AI

Abeba Birhane, William Isaac, Vinodkumar Prabhakaran, Mark Díaz, Madeleine Clare Elish, Iason Gabriel, Shakir Mohamed

专题命中 安全评测 :trustworthy(abstract);分类 cs.CY

Comments To appear in the proceeding of EAAMO 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.05300 2022-09-13 cs.LG cs.DC 57%

An Evaluation of Low Overhead Time Series Preprocessing Techniques for Downstream Machine Learning

Matthew L. Weiss, Joseph McDonald, David Bestor, Charles Yee, Daniel Edelman, Michael Jones, Andrew Prout, Andrew Bowne, Lindsey McEvoy, Vijay Gadepally, Siddharth Samsi

专题命中 安全评测 :alignment(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.09953 2022-08-23 stat.ML cs.LG stat.AP stat.ME 57%

Do-AIQ: A Design-of-Experiment Approach to Quality Evaluation of AI Mislabel Detection Algorithm

J. Lian, K. Choi, B. Veeramani, A. Hu, L. Freeman, E. Bowen, X. Deng

专题命中 安全评测 :trustworthy(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.07289 2022-08-16 cs.LG 57%

A Tool for Neural Network Global Robustness Certification and Training

Zhilu Wang, Yixuan Wang, Feisi Fu, Ruochen Jiao, Chao Huang, Wenchao Li, Qi Zhu

专题命中 安全评测 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.04273 2022-08-15 cs.AI 57%

Improving performance in multi-objective decision-making in Bottles environments with soft maximin approaches

Benjamin J Smith, Robert Klassert, Roland Pihlakas

专题命中 安全评测 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.04343 2022-08-12 cs.LG cs.LO 57%

EFI: A Toolbox for Feature Importance Fusion and Interpretation in Python

Aayush Kumar, Jimiama Mafeni Mase, Divish Rengasamy, Benjamin Rothwell, Mercedes Torres Torres, David A. Winkler, Grazziela P. Figueredo

专题命中 安全评测 :trustworthy(abstract);分类 cs.LG

Comments 16 pages, 5 tables, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.03096 2022-08-08 cs.LO cs.AI cs.PL 57%

Tools and Methodologies for Verifying Answer Set Programs

Zach Hansen

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

Comments In Proceedings ICLP 2022, arXiv:2208.02685

Journal ref EPTCS 364, 2022, pp. 211-216

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.01305 2022-08-03 cs.CY 57%

Humble Machines: Attending to the Underappreciated Costs of Misplaced Distrust

Bran Knowles, Jason D'Cruz, John T. Richards, Kush R. Varshney

专题命中 安全评测 :trustworthy(abstract);分类 cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2106.14956 2022-07-27 math.OC cs.DC cs.LG stat.ML 57%

Robust Distributed Optimization With Randomly Corrupted Gradients

Berkay Turan, Cesar A. Uribe, Hoi-To Wai, Mahnoosh Alizadeh

专题命中 安全评测 :trustworthy(abstract);分类 cs.LG

Comments 21 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2103.10689 2022-07-18 cs.LG 57%

Interpretable Deep Learning: Interpretation, Interpretability, Trustworthiness, and Beyond

Xuhong Li, Haoyi Xiong, Xingjian Li, Xuanyu Wu, Xiao Zhang, Ji Liu, Jiang Bian, Dejing Dou

专题命中 安全评测 :trustworthy(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.05796 2022-07-14 cs.LG 57%

Estimating Test Performance for AI Medical Devices under Distribution Shift with Conformal Prediction

Charles Lu, Syed Rakin Ahmed, Praveer Singh, Jayashree Kalpathy-Cramer

专题命中 安全评测 :safety(abstract);分类 cs.LG

Comments Principles of Distribution Shift (PODS) Workshop at ICML 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.12749 2022-07-04 cs.AI cs.HC 57%

A Human-Centric Assessment Framework for AI

Sascha Saralajew, Ammar Shaker, Zhao Xu, Kiril Gashteovski, Bhushan Kotnis, Wiem Ben Rim, Jürgen Quittek, Carolin Lawrence

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

Comments Accepted as submission to ICML 2022 Workshop on Human-Machine Collaboration and Teaming

详情

展开后加载摘要…

URL PDF HTML 收藏
2009.01492 2022-07-04 cs.LG stat.ML 57%

Explainable Empirical Risk Minimization

L. Zhang, G. Karakasidis, A. Odnoblyudova, L. Dogruel, A. Jung

专题命中 安全评测 :trustworthy(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.12685 2022-06-28 cs.CV cs.CR cs.LG 57%

Defense against adversarial attacks on deep convolutional neural networks through nonlocal denoising

Sandhya Aneja, Nagender Aneja, Pg Emeroylariffion Abas, Abdul Ghani Naim

专题命中 安全评测 :safety(abstract);分类 cs.LG

Journal ref IAES International Journal of Artificial Intelligence, Vol. 11, No. 3, September 2022, pp. 961~968, ISSN: 2252-8938

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.10063 2022-06-24 cs.RO cs.AI 57%

Resilient robot teams: a review integrating decentralised control, change-detection, and learning

David M. Bossens, Sarvapali Ramchurn, Danesh Tarapore

专题命中 安全评测 :safety(abstract);分类 cs.AI

Comments Accepted for Current Robotics Reports

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.09880 2022-06-22 cs.LG cs.CV 57%

Breaking Down Out-of-Distribution Detection: Many Methods Based on OOD Training Data Estimate a Combination of the Same Core Quantities

Julian Bitterwolf, Alexander Meinke, Maximilian Augustin, Matthias Hein

专题命中 安全评测 :trustworthy(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.09360 2022-06-22 cs.AI 57%

Modeling Transformative AI Risks (MTAIR) Project -- Summary Report

Sam Clarke, Ben Cottier, Aryeh Englander, Daniel Eth, David Manheim, Samuel Dylan Martin, Issa Rice

专题命中 安全评测 :safety(abstract);分类 cs.AI

Comments Chapters were written by authors independently. All authors are listed alphabetically

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.08304 2022-06-17 cs.CV cs.CR cs.LG eess.IV 57%

Adversarial Patch Attacks and Defences in Vision-Based Tasks: A Survey

Abhijith Sharma, Yijun Bian, Phil Munz, Apurva Narayan

专题命中 安全评测 :safety(abstract);分类 cs.LG

Comments A. Sharma and Y. Bian share equal contribution

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.08111 2022-06-17 cs.LG cs.CR math.OC stat.ML 57%

On Private Online Convex Optimization: Optimal Algorithms in $\ell_p$-Geometry and High Dimensional Contextual Bandits

Yuxuan Han, Zhicong Liang, Zhipeng Liang, Yang Wang, Yuan Yao, Jiheng Zhang

专题命中 安全评测 :trustworthy(abstract);分类 cs.LG

Comments This is the extended version of the paper appeared in the 39th International Conference on Machine Learning (ICML 2022): Optimal Private Streaming SCO in $\ell_p$-geometry with Applications in High Dimensional Online Decision Making

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.02051 2022-06-17 cs.AR cs.AI 57%

Fast and Accurate Error Simulation for CNNs against Soft Errors

Cristiana Bolchini, Luca Cassano, Antonio Miele, Alessandro Toschi

专题命中 安全评测 :safety(abstract);分类 cs.AI

Comments Accepted for publication in IEEE Transactions on Computers

详情

展开后加载摘要…

URL PDF HTML 收藏
2202.01136 2022-06-09 cs.LG cs.CV stat.ML 57%

Probabilistically Robust Learning: Balancing Average- and Worst-case Performance

Alexander Robey, Luiz F. O. Chamon, George J. Pappas, Hamed Hassani

专题命中 安全评测 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2007.10090 2022-06-03 cs.AI cs.LO 57%

Meet MASKS: A novel Multi-Classifier's verification approach

Amirhoshang Hoseinpour Dehkordi, Majid Alizadeh, Ali Movaghar

专题命中 安全评测 :safety(abstract);分类 cs.AI

Comments 34 pages, 12 figures, 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2106.06196 2022-05-31 cs.LG 57%

CausalAdv: Adversarial Robustness through the Lens of Causality

Yonggang Zhang, Mingming Gong, Tongliang Liu, Gang Niu, Xinmei Tian, Bo Han, Bernhard Schölkopf, Kun Zhang

专题命中 安全评测 :alignment(abstract);分类 cs.LG

Comments ICLR2022, 20 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.03673 2022-05-25 cs.LG eess.SP 57%

Towards Practical Physics-Informed ML Design and Evaluation for Power Grid

Shimiao Li, Amritanshu Pandey, Larry Pileggi

专题命中 安全评测 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.11487 2022-05-24 cs.CV cs.LG 57%

Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding

Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily Denton, Seyed Kamyar Seyed Ghasemipour, Burcu Karagol Ayan, S. Sara Mahdavi, Rapha Gontijo Lopes, Tim Salimans, Jonathan Ho, David J Fleet, Mohammad Norouzi

专题命中 安全评测 :alignment(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.06838 2022-05-18 cs.CL 57%

ePiC: Employing Proverbs in Context as a Benchmark for Abstract Language Understanding

Sayan Ghosh, Shashank Srivastava

专题命中 安全评测 :alignment(abstract);分类 cs.CL

Comments ACL 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
1911.11132 2022-05-17 cs.CV cs.LG 57%

Scaling Out-of-Distribution Detection for Real-World Settings

Dan Hendrycks, Steven Basart, Mantas Mazeika, Andy Zou, Joe Kwon, Mohammadreza Mostajabi, Jacob Steinhardt, Dawn Song

专题命中 安全评测 :safety(abstract);分类 cs.LG

Comments ICML 2022; The Species dataset and code are available at https://github.com/hendrycks/anomaly-seg

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.04678 2022-05-11 cs.LG stat.ML 57%

Real-time Forecasting of Time Series in Financial Markets Using Sequentially Trained Many-to-one LSTMs

Kelum Gajamannage, Yonggi Park

专题命中 安全评测 :trustworthy(abstract);分类 cs.LG

Comments 20 pages, 7 figures, submitted to Expert Systems with Applications Journal

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.14885 2022-05-09 cs.LG 57%

Out-of-Distribution Detection for Medical Applications: Guidelines for Practical Evaluation

Karina Zadorozhny, Patrick Thoral, Paul Elbers, Giovanni Cinà

专题命中 安全评测 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏