Metric Tools for Sensitivity Analysis with Applications to Neural Networks
专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG
Comments 15 pages
AI 大模型
大模型对齐、安全、越狱、红队、提示注入和可信评测。
专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG
Comments 15 pages
专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.AI
Comments 27 pages including supplemental information
专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG
专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG
Comments To appear at NeurIPS 2022
Journal ref https://proceedings.neurips.cc/paper_files/paper/2022/hash/867c06823281e506e8059f5c13a57f75-Abstract-Conference.html
专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.LG
专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.AI
Comments 35 pages, 15 figures; added GPT-4-base model results and discussion
专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG
Comments 15 pages, 3 figures, accepted for publication in the IEEE Transactions on Artificial Intelligence
Journal ref IEEE Transactions on Artificial Intelligence, 2023
专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG
Comments To appear in Proceedings of the 9th International Conference on Information Systems Security and Privacy
Journal ref Proceedings of the 9th International Conference on Information Systems Security and Privacy - ICISSP, pp. 339-348, 2023 , Lisbon, Portugal
专题命中 安全评测 :trustworthy(abstract);分类 cs.CY、cs.LG
Comments Accepted as a full paper (Best Paper nominee) at LAK 2023: The 13th International Learning Analytics and Knowledge Conference, March 13-17, 2023, Arlington, Texas, USA
专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG
专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG
Comments 19 pages, 5 tables, 7 figures, Annals of Telecommunications journal
专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG
专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.LG
Comments Accepted to ICLR 2023, final revision. https://openreview.net/forum?id=O-G91-4cMdv
专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG
Comments \c{opyright} Bhattacharya et al, 2023. This is the author's version of the work. It is posted here for your personal use. Not for redistribution. Copyright is held by the owner/author(s). Publication rights licensed to ACM. The definitive version was published in ACM IUI '23: 28th International Conference on Intelligent User Interfaces Proceedings, https://doi.org/10.1145/3581641.3584075
专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI
Comments 5 pages, 5 figures, 1 table, presented at AAAI 2023 conference for the AIAFS workshop
专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.CY
专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG
Comments Accepted for publication in IEEE VR 2023 conference. arXiv admin note: substantial text overlap with arXiv:2302.01985
专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG
Comments Accepted for publication in IUI 2023
专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG
专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG
专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.LG
Comments ICLR 2023 (accepted as Oral presentation)
专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG
Comments survey paper on holistic adversarial robustness for deep learning; published at AAAI 2023 Senior Member Presentation Track
专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG
Journal ref Frontiers in Artificial Intelligence, September 2022
专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG
Comments supplementary information in the main pdf
专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG
Comments 26 pages, 17 figures, 8 algorithms
Journal ref Nature Computational Science 2, 711 (2022)
专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG
专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG
Comments 23 pages (8 pages of references)
专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG
Comments 11 pages, 5 figures, Accepted at Reinforcement Learning for Real Life (RL4RealLife) Workshop at NeurIPS 2022
专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG
Comments 29 pages, 7 figures, 2 tables. IEEE Open Journal of the Communications Society (2022)
专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.LG
Comments EMNLP 2022 Camera-Ready (captions fixed)