Mimicking How Humans Interpret Out-of-Context Sentences Through Controlled Toxicity Decoding
模仿人类如何解读无上下文句子通过受控毒性解码
机构 * Department of Computer Science(计算机科学系)
AI总结 本文通过生成多样化的无上下文句子解读,模拟人类对不同毒性水平内容的感知,通过控制生成解读中的毒性来提高对人类写作解读的对齐性。
Comments Short paper; accepted at TrustNLP @ NAACL 2025