CommentsThis work has been accepted for publication at the IEEE Conference on Secure and Trustworthy Machine Learning (SaTML). The final version will be available on IEEE Xplore
Yuyan Bu, Xiaohao Liu, ZhaoXing Ren, Yaodong Yang, Juntao Dai
机构
*
Beijing Academy of Artificial Intelligence(北京人工智能研究院)
;
National University of Singapore(新加坡国立大学)
;
Institute for Artificial Intelligence, Peking University(北京大学人工智能研究院)
机构
*
School of Data Science, The Chinese University of Hong Kong, Shenzhen (CUHK-Shenzhen)(数据科学学院,香港中文大学(深圳))
;
Faculty of Information Technology, Monash University(信息技术学院,莫纳什大学)
;
Department of Dermatology, Tianjin Institute of Integrative Dermatology, Tianjin Academy of Traditional Chinese Medicine Affiliated Hospital(皮肤科,天津整合皮肤科研究所,天津中医研究院附属医院)
;
Department of Dermatology, The First Affiliated Hospital, Shantou University Medical College(皮肤科,汕头大学医学院第一附属医院)
;
Department of Dermatology, Beijing AnZhen Hospital, Capital Medical University(皮肤科,北京安贞医院,首都医科大学)
;
School of Computing and Data Science, The University of Hong Kong(计算与数据科学学院,香港大学)
;
Institute of Automation, Chinese Academy of Sciences(自动化研究所,中国科学院)
;
Department of Dermatology, Beijing Aerospace General Hospital(皮肤科,北京航天总医院)
机构
*
School of Physics, Peking University(物理系,北京大学)
;
School of Electronics Engineering and Computer Science, Peking University(电子工程与计算机科学系,北京大学)
;
Center for High Energy Physics, Peking University(高能物理中心,北京大学)
Intra-Fairness Dynamics: The Bias Spillover Effect in Targeted LLM Alignment
内在公平动态:定向大语言模型对齐中的偏见溢出效应
Eva Paraschou, Line Harder Clemmensen, Sneha Das
机构
*
Department of Applied Mathematics and Computer Science, Technical University of Denmark(应用数学与计算机科学系,丹麦技术大学)
;
Department of Mathematical Sciences, University of Copenhagen(数学科学系,哥本哈根大学)