Confident, Calibrated, or Complicit: Probing the Trade-offs between Safety Alignment and Ideological Bias in Language Models in Detecting Hate Speech
机构 * School of Physics, Mathematics and Computing(物理、数学与计算学系) ; Network Analysis and Social Influence Modelling (NASIM) Lab(网络分析与社会影响建模实验室) ; The University of Western Australia(西澳大学)
专题命中 AI治理与伦理 :alignment(title,abstract);safety(title,abstract);分类 cs.CL、cs.AI