The Hidden Space of Safety: Understanding Preference-Tuned LLMs in Multilingual context
机构 * LG Toronto AI Research lab(LG多伦多人工智能研究实验室)
专题命中 其他推理 :reasoning(abstract);分类 cs.CL
Comments 14 pages, 11 Figures, 2 Tables, currently under review at ACL 2025