发表机构
University of Louisville; Denison University(路易斯维尔大学; 丹尼森大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对大语言模型自动事实核查难题,DeLIVeR框架将证据检索设为强化探索任务,利用规划器大语言模型分解声明获问题集遍历知识图谱找证据,经GRPO优化策略,实验表明其显著优于基线,有效弥合推理差距,提供检测路径。
AI 中文摘要
由于传统检索系统中的“查询脆性”,自动事实核查对大语言模型(LLMs)来说仍然是一个挑战。我们提出了DeLIVeR(用于基于信息的真实性识别的分解学习)框架,将证据检索视为强化策略探索任务。它利用规划器大语言模型将复杂声明分解为目标问题集,用于遍历结构化知识图谱以获取高精度证据。通过组相对策略优化(GRPO)和优先考虑结构多样性及判定准确性的奖励系统来优化规划器策略。在LIAR、FEVER和PolitiFact上的评估表明,DeLIVeR显著优于现有基线。使用Qwen2.5 - 7B,该框架分别达到了83.73、84.57和79.70的峰值F1分数,比HippoRAG2提高了10 - 15%。通过转向强化问题规划策略,DeLIVeR有效弥合了多跳推理差距,为可验证的错误信息检测提供了可审计、透明的路径。
英文摘要
Automated fact-checking remains a challenge for Large Language Models (LLMs) due to "query brittleness" in traditional retrieval systems. We propose DeLIVeR (Decomposed Learning for Information-grounded Veracity Recognition), a framework that treats evidence retrieval as a reinforced strategic exploration task. DeLIVeR utilizes a Planner LLM to decompose complex claims into targeted question sets, which are used to traverse structured Knowledge Graphs (KGs) for high-precision evidence. We optimize the Planner's policy using Group Relative Policy Optimization (GRPO) with a reward system prioritizing structural diversity and verdict accuracy. Our evaluation on LIAR, FEVER, and PolitiFact shows that DeLIVeR significantly outperforms state-of-the-art baselines. Using Qwen2.5-7B, our framework achieved peak F1-scores of 83.73, 84.57, and 79.70 respectively, representing a 10-15% improvement over HippoRAG2. By shifting to a reinforced question-planning strategy, DeLIVeR effectively bridges multi-hop reasoning gaps and provides an auditable, transparent path for verifiable misinformation detection.
CommentsAccepted to 7th International Conference on Deep Learning Theory and Applications (DeLTA 2026)