arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 9380 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 9380 篇

2012.00419 2021-06-18 cs.LG cs.CY 62%

Machine Learning Systems in the IoT: Trustworthiness Trade-offs for Edge Intelligence

Wiebke Toussaint, Aaron Yi Ding

专题命中 安全评测 :trustworthy(abstract);分类 cs.CY、cs.LG

Comments In Proceedings of the Second International Conference on Cognitive Machine Intelligence (CogMI 2020)

详情

展开后加载摘要…

URL PDF HTML 收藏
2106.03379 2021-06-08 cs.CL cs.AI 62%

LAWDR: Language-Agnostic Weighted Document Representations from Pre-trained Models

Hongyu Gong, Vishrav Chaudhary, Yuqing Tang, Francisco Guzmán

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2103.11804 2021-06-08 cs.LG cs.CL cs.IR 62%

Detection of fake news on CoViD-19 on Web Search Engines

V. Mazzeo, A. Rapisarda, G. Giuffrida

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2102.05884 2021-06-08 cs.LG cs.AI stat.ML 62%

OpinionRank: Extracting Ground Truth Labels from Unreliable Expert Opinions with Graph-Based Spectral Ranking

Glenn Dawson, Robi Polikar

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

Comments 8 pages, 5 figures, accepted at IJCNN 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2103.12983 2021-06-03 cs.AI cs.LG 62%

Counterfactual Explanation with Multi-Agent Reinforcement Learning for Drug Target Prediction

Tri Minh Nguyen, Thomas P Quinn, Thin Nguyen, Truyen Tran

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2105.10026 2021-05-24 cs.CL cs.AI cs.CV 62%

Improving Generation and Evaluation of Visual Stories via Semantic Consistency

Adyasha Maharana, Darryl Hannan, Mohit Bansal

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

Comments NAACL 2021 (16 pages)

详情

展开后加载摘要…

URL PDF HTML 收藏
2105.06370 2021-05-14 cs.LG cs.CY 62%

Providing Assurance and Scrutability on Shared Data and Machine Learning Models with Verifiable Credentials

Iain Barclay, Alun Preece, Ian Taylor, Swapna K. Radha, Jarek Nabrzyski

专题命中 安全评测 :trustworthy(abstract);分类 cs.CY、cs.LG

Comments This is the submitted, pre-peer reviewed version of this paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2101.02494 2021-03-15 cs.LG cs.AI cs.SE 62%

Corner case data description and detection

Tinghui Ouyang, Vicent Sant Marco, Yoshinao Isobe, Hideki Asoh, Yutaka Oiwa, Yoshiki Seo

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2102.00714 2021-02-09 cs.LG cs.AI 62%

NeoRL: A Near Real-World Benchmark for Offline Reinforcement Learning

Rongjun Qin, Songyi Gao, Xingyuan Zhang, Zhen Xu, Shengkai Huang, Zewen Li, Weinan Zhang, Yang Yu

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2101.07685 2021-01-29 cs.LG cs.AI 62%

GLocalX -- From Local to Global Explanations of Black Box AI Models

Mattia Setzu, Riccardo Guidotti, Anna Monreale, Franco Turini, Dino Pedreschi, Fosca Giannotti

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

Comments 27 pages, 2 figures, submitted to "Special Issue on: Explainable AI (XAI) for Web-based Information Processing"

Journal ref Journal of Artificial Intelligence, Volume 294, May 2021, 103457

详情

展开后加载摘要…

URL PDF HTML 收藏
2010.07487 2021-01-21 cs.AI cs.CY 62%

Formalizing Trust in Artificial Intelligence: Prerequisites, Causes and Goals of Human Trust in AI

Alon Jacovi, Ana Marasović, Tim Miller, Yoav Goldberg

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.CY

Comments Accepted to ACM FAccT 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2012.11769 2020-12-23 cs.LG cs.AI 62%

Self-Progressing Robust Training

Minhao Cheng, Pin-Yu Chen, Sijia Liu, Shiyu Chang, Cho-Jui Hsieh, Payel Das

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

Comments Accepted in AAAI2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2011.09903 2020-11-20 cs.LG cs.AI 62%

Impact of Accuracy on Model Interpretations

Brian Liu, Madeleine Udell

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

Comments 7 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2011.04424 2020-11-11 cs.LG cs.AI cs.RO physics.optics 62%

Playing optical tweezers with deep reinforcement learning: in virtual, physical and augmented environments

Matthew Praeger, Yunhui Xie, James A. Grant-Jacob, Robert W. Eason, Ben Mills

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
1909.09962 2020-11-03 cs.CL cs.LG 62%

Adapting Language Models for Non-Parallel Author-Stylized Rewriting

Bakhtiyar Syed, Gaurav Verma, Balaji Vasan Srinivasan, Anandhavelu Natarajan, Vasudeva Varma

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.LG

Comments Accepted for publication in Main Technical Track at AAAI 20

详情

展开后加载摘要…

URL PDF HTML 收藏
2006.09526 2020-10-28 cs.CL cs.LG 62%

Cross-lingual Retrieval for Iterative Self-Supervised Training

Chau Tran, Yuqing Tang, Xian Li, Jiatao Gu

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.LG

Journal ref NeurIPS 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2004.14025 2020-10-08 cs.AI cs.CL 62%

Multi-View Attention Network for Visual Dialog

Sungjin Park, Taesun Whang, Yeochan Yoon, Heuiseok Lim

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2003.09761 2020-07-15 cs.CY cs.LG physics.soc-ph stat.ML 62%

Smarter Parking: Using AI to Identify Parking Inefficiencies in Vancouver

Devon Graham, Satish Kumar Sarraf, Taylor Lundy, Ali MohammadMehr, Sara Uppal, Tae Yoon Lee, Hedayat Zarkoob, Scott Duke Kominers, Kevin Leyton-Brown

专题命中 安全评测 :safety(abstract);分类 cs.CY、cs.LG

Comments All the authors contributed equally. This paper is an outcome of https://www.cs.ubc.ca/~kevinlb/teaching/cs532l%20-%202018-19/index.html. To be submitted to a journal in transportation or urban planning

详情

展开后加载摘要…

URL PDF HTML 收藏
1802.05027 2020-07-15 cs.LG cs.AI stat.ML 62%

Not to Cry Wolf: Distantly Supervised Multitask Learning in Critical Care

Patrick Schwab, Emanuela Keller, Carl Muroi, David J. Mack, Christian Strässle, Walter Karlen

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

Journal ref Proceedings of the 35th International Conference on Machine Learning, PMLR 80:4518-4527, 2018

详情

展开后加载摘要…

URL PDF HTML 收藏
2007.00753 2020-07-07 cs.LG cs.AI stat.ML 62%

Opportunities and Challenges in Deep Learning Adversarial Robustness: A Survey

Samuel Henrique Silva, Peyman Najafirad

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG

Comments 20 pages, 9 figures, submited to IEEE Transactions on Knowledge and Data Engineering

详情

展开后加载摘要…

URL PDF HTML 收藏
2006.11371 2020-06-24 cs.CV cs.AI cs.LG 62%

Opportunities and Challenges in Explainable Artificial Intelligence (XAI): A Survey

Arun Das, Paul Rad

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

Comments 24 pages, 20 figures, survey paper, submitting to IEEE

详情

展开后加载摘要…

URL PDF HTML 收藏
2006.10609 2020-06-19 cs.LG cs.AI stat.ML 62%

The Clever Hans Effect in Anomaly Detection

Jacob Kauffmann, Lukas Ruff, Grégoire Montavon, Klaus-Robert Müller

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

Comments 17 pages, preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
1910.12908 2020-03-10 cs.RO cs.AI cs.LG 62%

Certified Adversarial Robustness for Deep Reinforcement Learning

Björn Lütjens, Michael Everett, Jonathan P. How

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG

Comments Published at Conference on Robot Learning (CoRL) 2019; (v2) contains minor updates to related works; (v3) acknowledged AWS

Journal ref Proceedings of Machine Learning Research (PMLR) Vol. 100, 2019

详情

展开后加载摘要…

URL PDF HTML 收藏
1912.05100 2019-12-12 cs.LG cs.AI stat.ML 62%

Explainability Fact Sheets: A Framework for Systematic Assessment of Explainable Approaches

Kacper Sokol, Peter Flach

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG

Comments Conference on Fairness, Accountability, and Transparency (FAT* '20), January 27-30, 2020, Barcelona, Spain

详情

展开后加载摘要…

URL PDF HTML 收藏
1911.11253 2019-11-27 cs.LG cs.AI stat.ML 62%

Playing it Safe: Adversarial Robustness with an Abstain Option

Cassidy Laidlaw, Soheil Feizi

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
1911.10735 2019-11-26 cs.LG cs.AI cs.NE cs.PF 62%

CAMUS: A Framework to Build Formal Specifications for Deep Perception Systems Using Simulators

Julien Girard-Satabin, Guillaume Charpiat, Zakaria Chihani, Marc Schoenauer

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
1911.01275 2019-11-05 cs.CY cs.LG cs.SI 62%

Using Arabic Tweets to Understand Drug Selling Behaviors

Wesam Alruwaili, Bradley Protano, Tejasvi Sirigiriraju, Hamed Alhoori

专题命中 安全评测 :safety(abstract);分类 cs.CY、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
1905.09957 2019-10-29 cs.LG cs.AI cs.CR stat.ML 62%

Robust Attribution Regularization

Jiefeng Chen, Xi Wu, Vaibhav Rastogi, Yingyu Liang, Somesh Jha

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
1909.04436 2019-09-11 cs.LG cs.AI stat.AP stat.ML 62%

The Prevalence of Errors in Machine Learning Experiments

Martin Shepperd, Yuchen Guo, Ning Li, Mahir Arzoky, Andrea Capiluppi, Steve Counsell, Giuseppe Destefanis, Stephen Swift, Allan Tucker, Leila Yousefi

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

Comments 20th International Conference on Intelligent Data Engineering and Automated Learning (IDEAL), 14--16 November 2019

详情

展开后加载摘要…

URL PDF HTML 收藏
1908.02624 2019-08-08 cs.CY cs.AI 62%

A 20-Year Community Roadmap for Artificial Intelligence Research in the US

Yolanda Gil, Bart Selman

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.CY

Comments A Computing Community Consortium (CCC) workshop report, 109 pages

详情

展开后加载摘要…

URL PDF HTML 收藏