arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.34098cs.LG

超越正确性:评估跨表迁移中的语义知识

Beyond Correctness: Evaluating Semantic Knowledge in Cross-Table Transfer

Seokyong Sheem, Hochang Lee, Suyeong Lee, Daekyum Kim

首次发表
浏览论文内容

中文总结 AI 辅助

本研究揭示语义消融评估可能误导预测收益结论,提出区分内容敏感性与预测效用,并倡导根据研究问题选择对照的评估原则。

中文摘要 AI 辅助

语义知识越来越多地被用于弥合表格学习中的异构模式,但该知识实际上在多大程度上改善了预测?表格学习的研究通常通过语义消融来回答这个问题,即修改或抑制所提供的语义知识。我们表明,这些消融可能导致关于预测收益的误导性结论:在改变语义下的较差表现可能被当作证据,认为预期知识是有益的。在真实和受控实验中,改变语义内容可以产生较大的性能差异,即使模型在最初拥有该语义知识时几乎没有获得预测收益。为了区分这些效应,我们区分了两个量:内容敏感性和预测效用。内容敏感性衡量当语义内容被改变时性能的变化,而预测效用衡量预期语义知识相对于没有该知识的适当参考的收益。这一区分促使了一个评估框架,其中对照的选择取决于所提出的问题:改变的对照评估对语义内容的敏感性,而声称语义知识改善预测则需要适当的参考。即便如此,预测效用并非固定不变;它随适当的参考而变化,并且当参考更容易从其他输入或标记示例中恢复所测试的知识时,预测效用会降低。在对九项研究的25个语义消融比较的有界审计中,18个明确的预测效用声明中只有1个与能清晰隔离所测试语义贡献的对照配对。总之,这些发现促成了一个简单的评估原则:语义消融对照应根据其旨在回答的问题来选择并解释。

英文摘要

Semantic knowledge is increasingly used to bridge heterogeneous schemas in tabular learning, but how much does that knowledge actually improve prediction? Studies in tabular learning commonly answer this question through semantic ablations that modify or suppress the supplied semantic knowledge. We show that these ablations can lead to misleading conclusions about predictive benefit: poor performance under altered semantics may be taken as evidence that the intended knowledge is beneficial. Across real and controlled experiments, altering semantic content can produce large performance differences even when the model gains little predictive benefit from having that semantic knowledge in the first place. To separate these effects, we distinguish two quantities: content sensitivity and predictive utility. Content sensitivity measures the change in performance when semantic content is altered, whereas predictive utility measures the benefit of the intended semantic knowledge relative to a suitable reference without that knowledge. This distinction motivates an evaluation framework in which the control is chosen according to the question being asked: altered controls assess sensitivity to semantic content, whereas claims that semantic knowledge improves prediction require a suitable reference. Even then, predictive utility is not fixed; it varies across suitable references and decreases when the reference can more easily recover the tested knowledge from other inputs or labeled examples. In a bounded audit of 25 semantic-ablation comparisons across nine studies, only one of 18 explicit predictive-utility claims is paired with a control that clearly isolates the tested semantic contribution. Together, these findings motivate a simple evaluation principle: semantic-ablation controls should be chosen and interpreted according to the question they are intended to answer.

发表机构

  • Korea University(高丽大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑