Detecting and Filtering Unsafe Training Data via Data Attribution with Denoised Representation
机构 * University of Michigan(密歇根大学) ; University of Southern California(南加州大学) ; University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
Comments 14 pages
高校专区
机构 * University of Michigan(密歇根大学) ; University of Southern California(南加州大学) ; University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
Comments 14 pages
机构 * University of Southern California(南加州大学)
Comments 25 pages, 7 figures, 10 tables, currently under peer review
机构 * MIT(麻省理工学院) ; Microsoft(微软公司) ; University of Southern California(南加州大学) ; Microsoft Research(微软研究院) ; The University of Texas at Austin(德克萨斯大学奥斯汀分校)
机构 * University of Southern California(南加州大学) ; Northwestern University(西北大学) ; Arizona State University(亚利桑那州立大学) ; Adobe Research(Adobe研究) ; Rice University(里士满大学)
Comments Accepted to Findings of the Association for Computational Linguistics (ACL 2025), Vienna, Austria
机构 * University of Southern California(南加州大学) ; Arizona State University(亚利桑那州立大学)
Comments EMNLP Findings 2025. The project is available at https://github.com/USC-FORTIS/NLP-ADBench