arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.20537cs.LGcs.AI

可靠表格问答:可靠性标注需要多少监督?

ReliableTableQA:How Much Supervision Does Reliability Annotation Need?

Huei-Chung Hu, Hsin-Tai Wu, Koyo Kobayashi

首次发表
浏览论文内容

中文总结 AI 辅助

研究提出ReliableTableQA框架训练LLM标注表格问答结果统计可靠性,贡献分类法、数据管道及监督量研究。发现小的模式分层SFT集效果显著,GRPO在SFT不足时才有用,重新界定可靠性标注为数据效率问题,明确强化微调效果。

中文摘要 AI 辅助

我们引入了可靠表格问答(ReliableTableQA)框架,用于训练语言模型(LLM)来标注表格问答结果的统计可靠性,关注的不是查询是否可回答,而是计算出的答案在统计上是否有意义。在实际企业分析中,语法正确的SQL查询可能因样本过小、置信区间过宽或过于混淆而无法支持行动,但现有系统仍会自信回答,我们将此失败量化为不可靠自信答案率(UCAR)。我们贡献了:(1)涵盖小样本聚合、多重比较膨胀和分布尾部不匹配等风险的十类可靠性分类法(R1 - R10);(2)一个程序优先的数据管道,通过公共零售模式上的上下文无关语法生成50,000个带有可靠性标签的训练示例,并进行模式分层的SFT/GRPO分割;(3)一项关于校准可靠性标注实际所需监督量的对照研究。我们发现,一个小的、模式分层的SFT集就足够了:200个示例将可靠性标志F1从0.61提高到0.98,解析率从0.52提高到1.00,使UCAR降至零,并产生一个能推广到未见零售领域的模型(在保留的H&M上Rel - F1为0.997)。与这个强大的SFT基线相比,通常认为必不可少的GRPO仅在SFT训练不足时有所帮助(在100个示例时,分布内和分布外的精确标志集匹配提高0.06 - 0.16),一旦SFT足够,GRPO就没有可测量的益处,我们在硬复合标志切片、严格精确匹配度量和分布外评估中都证实了这一零结果。我们的发现将可靠性标注重新构建为一个数据效率问题,并精确界定了强化微调何时有效、何时无效。

英文摘要

We introduce ReliableTableQA, a framework for training an LLM to annotate the statistical reliability of tabular QA results, not whether the query is answerable, but whether the computed answer is statistically meaningful. In real enterprise analytics, a syntactically correct SQL query can return a value that is based on too small a sample, has an excessively wide confidence interval, or is too confounded to support action. Existing systems answer confidently in all such cases, a failure we quantify as the Unreliable Confident Answer Rate (UCAR). We contribute (1) a ten-category reliability taxonomy (R1-R10) covering hazards such as small-sample aggregates, multiple-comparison inflation, and distribution-tail mismatch; (2) a program-first data pipeline that generates 50,000 reliability-labeled training examples from a context-free grammar over public retail schemas, with schema-stratified SFT/GRPO splits; and (3) a controlled study of how much supervision calibrated reliability annotation actually requires. We find that a small, schema-stratified SFT set is remarkably sufficient: 200 examples raise reliability-flag F1 from 0.61 to 0.98 and parse rate from 0.52 to 1.00, drive UCAR to zero, and yield a model that generalizes to an unseen retail domain (Rel-F1 0.997 on held-out H&M). Against this strong SFT baseline, GRPO, commonly assumed to be essential, helps only when SFT is under-trained (+0.06-0.16 exact-flag-set match at 100 examples, in- and out-of-distribution) and provides no measurable benefit once SFT is adequate, a null result we confirm across a hard compound-flag slice, a strict exact-match metric, and out-of-distribution evaluation. Our findings reframe reliability annotation as a data-efficiency problem and delineate precisely when reinforcement fine-tuning does and does not pay off.

发表机构

  • DOCOMO Innovations, Inc.(都科摩创新公司)
  • NTT DOCOMO, Inc.(日本电报电话公司都科摩)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑