arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.03675cs.CL

VetScore:带引用的兽医长文本问答的风险加权事实核查

VetScore: Risk-Weighted Fact Verification for Veterinary Long-Form QA with Citations

Ivan Kartáč, Jan Tovarys, Mateusz Lango, Ondřej Dušek

AI总结:

本研究提出VetScore多步骤评估方法,针对兽医长文本问答,通过评估生成主张的伤害潜力与对引用片段的忠实度计算风险调整分数,经实验验证其与兽医专家评估高度相关且具备可解释性。

AI中文摘要:

引用片段可用于提升生成输出的可靠性及其对引用来源的忠实度,这在人类和兽医医学等高风险领域尤为重要。但这并不能保证生成的主张与提供的片段完全一致。我们提出VetScore,一种针对兽医长文本问答的多步骤评估方法,旨在评估生成主张在多大程度上得到提供片段的支持,并根据每个主张的伤害潜力对该信息进行加权。VetScore首先将输出分段并分解为单个主张,然后针对每个主张的伤害潜力进行评分并评估其对来源片段的忠实度,最后计算整体风险调整后的分数。我们收集了专家标注的元评估数据集,使用多种评判模型对我们的方法进行评估,结果表明,即使使用小型评判模型,它也能与兽医专家取得高度相关性,同时提供跨多个维度的可解释性。

英文摘要:

Citation excerpts can be used to increase the reliability of generated outputs and their faithfulness to cited sources, which is especially important in high-stakes domains such as human and veterinary medicine. However, this does not guarantee that generated claims are faithful to the provided excerpts. We present VetScore, a multi-step evaluation method for veterinary long-form question answering, designed to assess how well are generated claims supported by the provided excerpts, weighing this information by each claim's harm potential. VetScore first segments the output and decomposes it into individual claims, then scores each claim with respect to its harm potential and evaluates its faithfulness to source excerpts, and finally calculates the overall risk-adjusted score. We collect an expert-annotated meta-evaluation dataset, evaluate our approach with a range of judge models, and show that it achieves high correlations with veterinary experts even with small judge models, while offering explainability across multiple dimensions.

↑