arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 9380 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 9380 篇

2203.12670 2022-03-25 cs.LG cs.AI cs.HC cs.NE cs.RO 62%

Competency Assessment for Autonomous Agents using Deep Generative Models

Aastha Acharya, Rebecca Russell, Nisar R. Ahmed

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.11147 2022-03-22 cs.CL cs.LG 62%

Teaching language models to support answers with verified quotes

Jacob Menick, Maja Trebacz, Vladimir Mikulik, John Aslanides, Francis Song, Martin Chadwick, Mia Glaese, Susannah Young, Lucy Campbell-Gillingham, Geoffrey Irving, Nat McAleese

专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.03712 2022-03-15 cs.CY cs.AI 62%

Trusted Data Forever: Is AI the Answer?

Emanuele Frontoni, Marina Paolanti, Tracey P. Lauriault, Michael Stiber, Luciana Duranti, Abdul-Mageed Muhammad

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.04386 2022-03-10 cs.LG cs.AI cs.IT eess.SP math.IT 62%

Model-free feature selection to facilitate automatic discovery of divergent subgroups in tabular data

Girmaw Abebe Tadesse, William Ogallo, Celia Cintas, Skyler Speakman

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.09271 2022-03-07 cs.CY cs.AI 62%

Rebuilding Trust: Queer in AI Approach to Artificial Intelligence Risk Management

Ashwin, William Agnew, Umut Pajaro, Hetvi Jethwani, Arjun Subramonian

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.CY

Comments We discovered that the manuscript unintentionally contains passages that are direct quotes from previous literature, but fails to properly address them as such

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.15035 2022-03-01 cs.LG cs.AI 62%

Focus! Rating XAI Methods and Finding Biases

Anna Arias-Duart, Ferran Parés, Dario Garcia-Gasulla, Victor Gimenez-Abalos

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.02992 2022-03-01 cs.CL cs.AI 62%

Towards More Robust Natural Language Understanding

Xinliang Frederick Zhang

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.AI

Comments Undergraduate Research Thesis, The Ohio State University

详情

展开后加载摘要…

URL PDF HTML 收藏
2202.05302 2022-02-14 cs.LG cs.AI 62%

Trust in AI: Interpretability is not necessary or sufficient, while black-box interaction is necessary and sufficient

Max W. Shen

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2202.03199 2022-02-08 cs.AI cs.LG cs.SC 62%

AI Research Associate for Early-Stage Scientific Discovery

Morad Behandish, John Maxwell, Johan de Kleer

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

Comments Paper #203

Journal ref AAAI-MLPS-2021: Association for the Advancement of Artificial Intelligence (AAAI) 2021 Spring Symposium on Combining Artificial Intelligence and Machine Learning with Physics Sciences (MLPS)

详情

展开后加载摘要…

URL PDF HTML 收藏
2202.03059 2022-02-08 cs.RO cs.AI cs.CV cs.LG 62%

Evaluation of Runtime Monitoring for UAV Emergency Landing

Joris Guerin, Kevin Delmas, Jérémie Guiochet

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG

Comments 7 pages, 4 figures, 1 table. To appear in the proceedings of 2022 IEEE International Conference on Robotics and Automation (ICRA)

详情

展开后加载摘要…

URL PDF HTML 收藏
2102.09677 2022-01-26 cs.LG cs.AI cs.NE cs.RO cs.SY eess.SY 62%

Training a Resilient Q-Network against Observational Interference

Chao-Han Huck Yang, I-Te Danny Hung, Yi Ouyang, Pin-Yu Chen

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG

Comments Accepted to AAAI 2022. 9 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2106.09296 2022-01-17 cs.LG cs.AI cs.NE cs.SD eess.AS 62%

Voice2Series: Reprogramming Acoustic Models for Time Series Classification

Chao-Han Huck Yang, Yun-Yun Tsai, Pin-Yu Chen

专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.LG

Comments Updated version with a correction. The full draft was submitted in Jan 2021. The Voice2Series project initially was launched in Sep 2020. Accepted to ICML 2021, 16 Pages

Journal ref Proceedings of the 38th International Conference on Machine Learning 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2101.01041 2022-01-03 math.OC cs.AI cs.LG cs.SY eess.SY 62%

Derivative-Free Policy Optimization for Linear Risk-Sensitive and Robust Control Design: Implicit Regularization and Sample Complexity

Kaiqing Zhang, Xiangyuan Zhang, Bin Hu, Tamer Başar

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2104.10726 2021-12-07 cs.CL cs.AI 62%

Learning Fine-grained Fact-Article Correspondence in Legal Cases

Jidong Ge, Yunyun huang, Xiaoyu Shen, Chuanyi Li, Wei Hu

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.AI

Comments Code and dataset are available at https://github.com/gjdnju/MLMN

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.06721 2021-12-02 cs.LG cs.AI stat.ML 62%

Causal Multi-Agent Reinforcement Learning: Review and Open Problems

St John Grimbly, Jonathan Shock, Arnu Pretorius

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG

Comments Accepted at Cooperative AI Workshop, NeurIPS 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2107.04457 2021-11-17 cs.RO cs.AI cs.LG physics.optics 62%

Aligning an optical interferometer with beam divergence control and continuous action space

Stepan Makarenko, Dmitry Sorokin, Alexander Ulanov, A. I. Lvovsky

专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.LG

Comments 12 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.02739 2021-11-08 cs.LG cs.AI cs.CV cs.RO 62%

A Step Towards Efficient Evaluation of Complex Perception Tasks in Simulation

Jonathan Sadeghi, Blaine Rogers, James Gunn, Thomas Saunders, Sina Samangooei, Puneet Kumar Dokania, John Redford

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG

Comments To appear in NeurIPS 2021 Workshop on Machine Learning for Autonomous Driving (ML4AD)

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.02400 2021-11-05 cs.LG cs.AI cs.CV eess.IV math.OC 62%

Deep AUC Maximization for Medical Image Classification: Challenges and Opportunities

Tianbao Yang

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

Comments Medical Imaging meets NeurIPS 2021 workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.02398 2021-11-05 eess.IV cs.AI cs.CV cs.LG 62%

Transparency of Deep Neural Networks for Medical Image Analysis: A Review of Interpretability Methods

Zohaib Salahuddin, Henry C Woodruff, Avishek Chatterjee, Philippe Lambin

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.15527 2021-11-01 cs.CL cs.AI 62%

Pre-training Co-evolutionary Protein Representation via A Pairwise Masked Language Model

Liang He, Shizhuo Zhang, Lijun Wu, Huanhuan Xia, Fusong Ju, He Zhang, Siyuan Liu, Yingce Xia, Jianwei Zhu, Pan Deng, Bin Shao, Tao Qin, Tie-Yan Liu

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.14222 2021-10-28 cs.LG cs.AI stat.ML 62%

Sample Selection for Fair and Robust Training

Yuji Roh, Kangwook Lee, Steven Euijong Whang, Changho Suh

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

Comments Accepted to 35th Conference on Neural Information Processing Systems (NeurIPS), 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.12588 2021-10-26 cs.LG cs.AI cs.SE 62%

QuantifyML: How Good is my Machine Learning Model?

Muhammad Usman, Divya Gopinath, Corina S. Păsăreanu

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG

Comments In Proceedings FMAS 2021, arXiv:2110.11527

Journal ref EPTCS 348, 2021, pp. 92-100

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.08956 2021-10-20 eess.SY cs.AI cs.LG cs.SY 62%

Improving Robustness of Reinforcement Learning for Power System Control with Adversarial Training

Alexander Pan, Yongkyun Lee, Huan Zhang, Yize Chen, Yuanyuan Shi

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG

Comments Published at 2021 ICML RL4RL Workshop; Submitted to 2022 PSCC

详情

展开后加载摘要…

URL PDF HTML 收藏
2011.08484 2021-09-29 cs.RO cs.AI cs.LG 62%

Combining Reinforcement Learning with Model Predictive Control for On-Ramp Merging

Joseph Lubars, Harsh Gupta, Sandeep Chinchali, Liyun Li, Adnan Raja, R. Srikant, Xinzhou Wu

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG

Comments 8 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2108.05170 2021-09-10 cs.CV cs.AI cs.LG 62%

ProAI: An Efficient Embedded AI Hardware for Automotive Applications -- a Benchmark Study

Sven Mantowsky, Falk Heuer, Syed Saqib Bukhari, Michael Keckeisen, Georg Schneider

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG

Comments Accepted by IEEE International Conference on Computer Vision (ICCV) 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2108.07380 2021-08-23 stat.ML cs.AI cs.LG econ.EM 62%

InfoGram and Admissible Machine Learning

Subhadeep Mukhopadhyay

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

Comments Keywords: Admissible machine learning; InfoGram; L-Features; Information-theory; ALFA-testing, Algorithmic risk management; Fairness; Interpretability; COREml; FINEml

详情

展开后加载摘要…

URL PDF HTML 收藏
2108.06159 2021-08-16 cs.CV cs.AI cs.LG 62%

Robustness testing of AI systems: A case study for traffic sign recognition

Christian Berghoff, Pavol Bielik, Matthias Neu, Petar Tsankov, Arndt von Twickel

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG

Comments 12 pages, 7 figures. The final publication is available at Springer via https://doi.org/10.1007/978-3-030-79150-6_21

Journal ref In: Maglogiannis I., Macintyre J., Iliadis L. (eds) Artificial Intelligence Applications and Innovations. AIAI 2021. IFIP Advances in Information and Communication Technology, vol 627. Springer, Cham

详情

展开后加载摘要…

URL PDF HTML 收藏
2108.01851 2021-08-05 cs.LG cs.AI cs.RO 62%

Risk Conditioned Neural Motion Planning

Xin Huang, Meng Feng, Ashkan Jasour, Guy Rosman, Brian Williams

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG

Comments Accepted at IROS'21. Author version with 7 pages, 5 figures, 2 tables, and 1 algorithm

详情

展开后加载摘要…

URL PDF HTML 收藏
2105.03371 2021-07-14 cs.DC cs.AI cs.DB cs.LG cs.MA 62%

The Synergy of Complex Event Processing and Tiny Machine Learning in Industrial IoT

Haoyu Ren, Darko Anicic, Thomas Runkler

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG

Comments Accepted by The 15th ACM International Conference on Distributed and Event-based Systems (DEBS) 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2101.11453 2021-06-23 cs.LG cs.AI cs.CV stat.ML 62%

Meta Adversarial Training against Universal Patches

Jan Hendrik Metzen, Nicole Finnie, Robin Hutmacher

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG

Comments Accepted by the ICML 2021 workshop on "A Blessing in Disguise: The Prospects and Perils of Adversarial Machine Learning"

详情

展开后加载摘要…

URL PDF HTML 收藏