The Unintended Trade-off of AI Alignment:Balancing Hallucination Mitigation and Safety in LLMs
AI对齐中的意外权衡:在LLMs中平衡幻觉缓解与安全
机构 * Applied Artificial Intelligence Initiative(应用人工智能倡议) ; Deakin University(德肯大学) ; School of Information Technology(信息科技学院)
专题命中 幻觉与事实性 :alignment(title,abstract);safety(title,abstract);分类 cs.CL
AI总结 本文提出了一种方法,通过解耦幻觉和拒绝特征,平衡LLMs中的真实性与安全性,减少因提高真实性而削弱安全对齐的问题。