发表机构
Aidentyx Inc.(艾登蒂克斯公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究闭环表格识别中LLM-as-a-judge的效用,发现其评判信号弱,无特定反馈也有严重损失,结构保留约束效果不佳,表明评估能力不意味着优化效用,迭代细化需确定性检测结构变化的验证信号。
AI 中文摘要
大语言模型作为评判器在闭环再生中广泛用于提供反馈和选择信号,但该用途尚未得到充分验证。我们在表格识别中进行研究,利用确定性TEDS评估提供可控测试平台,使用FinTabNet和OmniDocBench数据集。有三个发现:一是评判信号在两个数据集上都很弱;二是即使没有特定评判反馈也会出现严重损失,结构保留指令可降低严重损失率;三是结构保留约束减少了严重损失尾部但无改善。这些结果表明评估能力不意味着优化效用,迭代细化至少需要能确定性检测结构变化的验证信号,而非仅靠评判分数。
英文摘要
LLM-as-a-judge is widely used to provide feedback and selection signals in closedloop regeneration, but this use remains insufficiently validated. We study it in table recognition, where deterministic TEDS evaluation provides a controlled testbed, using FinTabNet and OmniDocBench. Three findings emerge. First, judge signals were weak on both datasets: scores frequently tied, rankings were not reproducible, and no tested judge score policy, whether selecting among candidates or accepting revisions under a conservative score margin, improved on the first output on both datasets. Iteration produced better candidates, but the judge recovered them at most partially on one dataset and not at all on the other. Second, severe losses occurred even without specific judge feedback, supporting target-preservation failure under unconstrained regeneration as a proximate mechanism. Third, a structure-preserving instruction reduced the severe-loss rate, significantly on FinTabNet and directionally on OmniDocBench, but produced no improvement, and in an exploratory 2x2 analysis this protection was not stably observed when judge feedback was retained. These results do not dispute the value of LLMs as evaluators, but show that the tested reference-free judge signals were too weak and unstable to drive candidate selection in this setup, and that evaluation-style evidence alone was insufficient to establish closed-loop optimization utility. Iterative refinement requires, at minimum, a verification signal that deterministically detects structural change, rather than judge scores alone.
Comments32 pages, 9 figures, appendix included