The Impact of Coreset Selection on Spurious Correlations and Group Robustness
机构 * Princeton University(普林斯顿大学) ; FAIR at Meta(Meta 的 FAIR 实验室)
专题命中 安全评测 :alignment(abstract);分类 cs.LG
Comments 10 pages, 9 additional pages for Appendix
AI 大模型
大模型对齐、安全、越狱、红队、提示注入和可信评测。
机构 * Princeton University(普林斯顿大学) ; FAIR at Meta(Meta 的 FAIR 实验室)
专题命中 安全评测 :alignment(abstract);分类 cs.LG
Comments 10 pages, 9 additional pages for Appendix
专题命中 安全评测 :safety(abstract)
专题命中 安全评测 :alignment(abstract)
Comments Under review
机构 * The Fu Foundation School of Engineering and Applied Science, Columbia University(哥伦比亚大学福基金会工程与应用科学学院) ; School of Nursing, Columbia University(哥伦比亚大学护理学院) ; Center for Home Care Policy & Research, VNS Health(VNS健康居家护理政策与研究中心)
专题命中 安全评测 :alignment(abstract)
Comments The Second Workshop on GenAI for Health at NeurIPS 2025
机构 * University of Surrey, UK(Surrey大学)
专题命中 安全评测 :alignment(abstract)
Comments Accepted at the NeurIPS 2024 Workshop on Audio Imagination; this version updates the project page link
机构 * M42, Abu Dhabi(阿布扎比M42)
专题命中 AI治理与伦理 :safety(abstract);分类 cs.CL
Comments Accepted to EMNLP Main 2025
专题命中 AI治理与伦理 :alignment(abstract)
Comments 10 pages, 1 figure, 2 tables
专题命中 其他安全 :safety(title,abstract);AI safety(title,abstract);分类 cs.AI、cs.CY
机构 * MARS ; Redwood Research
专题命中 其他安全 :safety(title,abstract);AI safety(title);分类 cs.AI、cs.CY
机构 * Universitat Politècnica de Catalunya(加泰罗尼亚理工大学)
专题命中 其他安全 :alignment(title,abstract);分类 cs.AI
机构 * Tencent YouTu Lab(腾讯优图实验室) ; East China University of Science and Technology(东华大学) ; Peking University(北京大学) ; Renmin University of China(中国人民大学) ; Shenzhen University(深圳大学) ; Hong Kong University of Science and Technology(香港科技大学)
专题命中 其他安全 :alignment(title,abstract)
Comments NeurIPS 2025 Spotlight. 13 Pages, 10 figures
机构 * Yuan Ze University(元智大学)
专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG
Comments Accepted by The 38th Conference of Open Innovations Association FRUCT, 2025
机构 * School of Computer Science and Informatics, University of Liverpool, UK(计算机科学与信息学学院,利物浦大学)
专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.AI
Comments Accepted at NeurIPS 2025 with minor revisions
机构 * Novus Technologies ; MIT(麻省理工学院) ; Harvard University(哈佛大学)
专题命中 其他安全 :alignment(abstract,comments);分类 cs.CL
Comments ICLR 2024 Workshop on Representational Alignment (Re-Align) Camera Ready
专题命中 其他安全 :alignment(abstract);分类 cs.LG
Comments NeurIPS 2025
专题命中 其他安全 :alignment(abstract);分类 cs.CL
Comments Link: https://github.com/BaileyWei/BMEmbed
Journal ref Findings of the Association for Computational Linguistics ACL 2025
机构 * Mongol AI(蒙古AI)
专题命中 其他安全 :alignment(abstract);分类 cs.CL
Comments 12 pages, 3 figures, 4 tables
专题命中 其他安全 :safety(abstract)
Comments At IEEE S&P 2026