发表机构
Huawei Technologies(华为技术有限公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对大模型幻觉问题,提出HARPO强化学习框架,通过幻觉感知奖励模型与选择性激活机制,在提升忠实性的同时增强创造性,实验验证了其有效性。
AI 中文摘要
大型语言模型(LLMs)容易生成幻觉内容,这损害了它们在知识密集型任务中的可靠性。为了在不牺牲创造性的情况下应对这一挑战,我们提出了HARPO,一个旨在同时优化忠实性和创造性的强化学习框架。HARPO引入了一个通过可验证反馈训练的幻觉感知生成奖励模型(HA-GRM),用于评估忠实性和写作质量。选择性激活机制(SAM)仅对HA-GRM判定为无幻觉的输出激活写作奖励,而数据课程则逐步将训练从创造性写作转向以幻觉为中心的任务。在RAGTruth上,我们基于Qwen3-4B的HA-GRM达到了78.08%的响应级F1分数,而监督微调基线为66.37%。在参数规模从1.7B到8B的Qwen2.5和Qwen3模型上的实验显示,忠实生成和写作质量均有所提升。在Qwen3-4B上,HARPO将MultiHopRAG上由HA-GRM判定的幻觉率从3.29%降至1.02%,同时将Arena-Hard-v2.0的创造性写作分数从16.95%提升至27.54%。
英文摘要
Large Language Models (LLMs) are prone to generating hallucinated content, which compromises their reliability in knowledge-intensive tasks. To address this challenge without sacrificing creativity, we propose HARPO, a reinforcement learning framework designed to jointly optimize faithfulness and creativity. HARPO incorporates a Hallucination-Aware Generative Reward Model (HA-GRM), trained via verifiable feedback, to assess both faithfulness and writing quality. A Selective Activation Mechanism (SAM) activates writing rewards only for outputs judged hallucination-free by HA-GRM, while a data curriculum progressively shifts training from creative writing to hallucination-centric tasks. On RAGTruth, our Qwen3-4B-based HA-GRM achieves a response-level F1 score of 78.08%, compared with 66.37% for the supervised fine-tuning baseline. Experiments on Qwen2.5 and Qwen3 models from 1.7B to 8B parameters show improvements in both faithful generation and writing quality. On Qwen3-4B, HARPO reduces the HA-GRM-judged hallucination rate on MultiHopRAG from 3.29% to 1.02%, while increasing the Arena-Hard-v2.0 creative-writing score from 16.95% to 27.54%.
Comments11 pages