XQDT:带反馈信号的可解释定量数据-文本对齐指标
XQDT: eXplainable and Quantitative Data-Text Alignment Metric with Feedback Signals
浏览论文内容
中文总结 AI 辅助
本文提出带反馈信号的可解释定量数据-文本对齐指标XQDT,通过微调语言模型识别数据-文本对的各类数据单元,在基准测试中性能优于LLM作为评判方法,可用于评估及下游修正优化。
中文摘要 AI 辅助
评估数据-文本对齐仍具挑战性:现有指标往往对分数的解释有限,而基于提示的大语言模型(LLM)作为评判方法成本高且不可靠。本文提出一种端到端可解释评估指标,通过微调语言模型来识别数据-文本对中被遗漏、多余、错误和正确的数据单元。这些局部判断被汇总为精确率、召回率和F1分数,既提供细粒度诊断反馈,又提供可解释的对齐质量度量。在基准测试中,本文微调的模型在错误预测方面优于LLM作为评判方法,且精确率、召回率和F1分数具有竞争力,同时与人类判断保持强相关性。除评估外,本文验证器的输出还为下游修正与优化提供有用的反馈信号,支持面向对齐的数据到文本和文本到数据的改进。代码和资源可在该https URL获取。
英文摘要
Evaluating data-text alignment remains challenging: existing metrics often provide limited explanations for the scores, while prompt-based LLM-as-Judge methods can be expensive and unreliable. We present an end-to-end explainable evaluation metric that fine-tunes a language model to identify omitted, extra, incorrect, and correct data units in a data-text pair. These local judgements are aggregated into precision, recall, and F1 scores, providing both fine-grained diagnostic feedback and an interpretable measure of alignment quality. Across benchmarks, our fine-tuned models outperform LLM-as-Judge methods in error prediction and achieve competitive precision, recall, and F1 scores, while maintaining strong correlation with human judgements. Beyond evaluation, our verifier outputs also provide useful feedback signals for downstream correction and refinement, supporting alignment-oriented improvement of data-to-text and text-to-data. Code and resources are available at https://github.com/guihuzhang/xqdt.
发表机构
- CNRS/LORIA(法国国家科学研究中心/洛里亚研究所)
- Université de Lorraine(洛林大学)
机构由 AI 辅助整理,请以论文原文为准。