TriQua:在事实性评估中协调粒度与上下文
TriQua: Reconciling Granularity and Context in Factuality Evaluation
浏览论文内容
中文总结 AI 辅助
TriQua框架解决LLM事实性评估的粒度与上下文权衡问题,提出自适应事实建模方式及TriQuaScore,在事实验证任务中表现优于现有框架且与人工评分高度一致。
中文摘要 AI 辅助
针对大语言模型(LLM)事实性评估的“先分解后验证”范式存在根本权衡:原子事实(即传递一个信息单元的单句)常缺失必要上下文,而更宽泛的陈述则缺乏精确评估所需的粒度。为解决这一问题,我们提出TriQua框架,该框架可根据事实的复杂度灵活建模:简单主张被提取为标准三元组,复杂主张则通过附加辅助上下文限定词表示为超关系事实。这种自适应结构在不牺牲原子性的前提下,保留了准确检索与验证所需的必要上下文。此外,TriQua的验证过程直接标注特定三元组和限定词内的具体错误,为错误检测提供细粒度可解释性。伴随该框架,我们提出TriQuaScore以量化这些结构化事实单元的事实性。实证评估显示,TriQuaScore与人工标注的事实性评分高度一致,TriQua实现了稳健的分解质量,且在基于证据的事实验证中优于现有基于分解的框架。
英文摘要
The "decompose-then-verify" paradigm for LLM factuality evaluation faces a fundamental trade-off: atomic facts, i.e., one sentence conveying one unit of information, often omit essential context, while broader statements lack the granularity needed for precise assessment. To address this, we introduce TriQua, a framework that flexibly models facts based on their complexity. Simple claims are extracted as standard triples, while complex claims are represented as hyperrelational facts by attaching auxiliary contextual qualifiers. This adaptive structure preserves the necessary context for accurate retrieval and verification without sacrificing atomicity. Furthermore, TriQua's verification process directly annotates concrete errors within specific triples and qualifiers, providing fine-grained explainability for error detection. Alongside the framework, we propose TriQuaScore to quantify the factuality of these structured fact units. Empirical evaluations show that TriQuaScore strongly aligns with human annotated factuality scores, TriQua achieves robust decomposition quality, and outperforms existing decomposition-based frameworks in evidence-based fact verification.
发表机构
- FZI Research Center for Information Technology(FZI信息技术研究中心)
- Karlsruhe Institute of Technology(卡尔斯鲁厄理工学院)
- Trier University(特里尔大学)
机构由 AI 辅助整理,请以论文原文为准。