Quantifying the Importance of Data Alignment in Downstream Model Performance
机构 * stanford(斯坦福大学)
专题命中 安全评测 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG
Journal ref ICLR DMLR Data-centric Machine Learning Research (2024), ICML DataWorld (2025)
AI 大模型
大模型对齐、安全、越狱、红队、提示注入和可信评测。
机构 * stanford(斯坦福大学)
专题命中 安全评测 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG
Journal ref ICLR DMLR Data-centric Machine Learning Research (2024), ICML DataWorld (2025)
机构 * School of Computing and Information System, the University of Melbourne, Australia(墨尔本大学计算与信息学院) ; School of Computing, FSE, Macquarie University, Australia(麦考瑞大学计算学院)
专题命中 安全评测 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG
Comments Accepted to ACL 2025 (Main Proceedings)
机构 * Oracle AI ; Indian Institute of Information Technology Ranchi(印度拉奇信息与技术学院) ; TD Securities(TD证券) ; Columbia University(哥伦比亚大学) ; Hanyang University(翰阳大学)
专题命中 安全评测 :safety(title);alignment(abstract);分类 cs.CL、cs.AI、cs.LG
Comments Published in the Proceedings of the 2025 Conference of the North American Chapter of the Association for Computational Linguistics (NAACL 2025), Industry Track, pages 558-582
机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) ; Amazon(亚马逊) ; Fidelity Investments(富达投资)
专题命中 安全评测 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.CY
Comments 21 pages, 11 figures, 7 tables
机构 * Warren E Hyde Middle School(沃伦·E·海德中学) ; University of Trento(特伦托大学)
专题命中 安全评测 :safety(title,abstract);分类 cs.CL、cs.AI、cs.CY
Comments Accepted to ACM FAccT 2025. To be presented in Athens, June 2025, and published in the conference proceedings. Preprint version; final version will appear in the ACM Digital Library
机构 * Unicom Data Intelligence(中国联通数据智能研究所) ; Data Science & Artificial Intelligence Research Institute(数据科学与人工智能研究院)
专题命中 安全评测 :safety(title,abstract);分类 cs.CL、cs.AI、cs.CY
Comments 21 pages, 13 figures, 4 tables
专题命中 安全评测 :safety(title,abstract);分类 cs.CL、cs.AI、cs.CY
专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI、cs.CY、cs.LG
专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI、cs.CY、cs.LG
Comments This work has been accepted for publication in AI and Ethics
专题命中 安全评测 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG
Comments 23 pages, 13 tables, 3 figures
专题命中 安全评测 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG
Comments Working in progress
专题命中 安全评测 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG
专题命中 安全评测 :alignment(title,abstract);trustworthy(abstract)
专题命中 安全评测 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG
Comments Accepted to ACL 2024
专题命中 安全评测 :safety(title,abstract);alignment(abstract)
Comments 18 pages, 3 figures, accepted for publication at the International Symposium on Leveraging Applications of Formal Methods (ISoLA 2024)
专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI、cs.CY、cs.LG
专题命中 安全评测 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG
Comments Dataset and models are available at https://github.com/jihaonew/MM-Instruct
专题命中 安全评测 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG
Comments 9 pages (excluding references), accepted to ACL 2024 Main Conference
专题命中 安全评测 :safety(title,abstract);分类 cs.CL、cs.AI、cs.LG
Comments Accepted at ACL 2024 Findings
专题命中 安全评测 :trustworthy(title,abstract);分类 cs.CL、cs.AI、cs.LG
Comments Accepted to ACL
专题命中 安全评测 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG
Comments v3 is the camera-ready version for AAAI-24 Workshop on Public Sector LLMs: Algorithmic and Sociotechnical Design. 7 pages of main content, 1 page of references, 3 pages of appendices, and 7 figures. Our full prompts are released in the repo: https://github.com/zowiezhang/HVAE
专题命中 安全评测 :safety(title,abstract);分类 cs.CL、cs.AI、cs.LG
Comments A preprint work. Benchmark link: https://github.com/whitzard-ai/jade-db. Website link: https://whitzard-ai.github.io/jade.html
专题命中 安全评测 :safety(title,abstract);分类 cs.CL、cs.AI、cs.LG
Comments To be published in GEM workshop. Conference on Empirical Methods in Natural Language Processing (EMNLP). 2023
专题命中 安全评测 :trustworthy(title,abstract);分类 cs.CL、cs.AI、cs.LG
专题命中 安全评测 :safety(title,abstract);分类 cs.CL、cs.AI、cs.CY
Comments main paper p.1-29, 5 figures, 2 tables
专题命中 安全评测 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.CY
Comments 20 pages, 5 figures
专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI、cs.CY、cs.LG
专题命中 安全评测 :safety(title,abstract);AI safety(abstract)
专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI、cs.CY、cs.LG
专题命中 安全评测 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG
Comments Full version (with appendix) of the accepted paper in 36th AAAI Conference on Artificial Intelligence 2022