发表机构
Technical University of Darmstadt; DFKI(达姆施塔特工业大学; 德国人工智能研究中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究表格级嵌入,通过扩展TEmBed进行系统评估,捕捉下游有效性所需属性。对TEmBed模型池实证研究发现无单一模型在所有任务中皆优,证明表格级嵌入质量不能仅用检索衡量。
AI 中文摘要
表格数据是主要的结构化数据形式,学习表格表示已成为核心研究方向。表格级嵌入尤其支撑着广泛应用,如表格检索、数据湖发现和表格分类。尽管其很重要,但对不同嵌入方法在各任务中的表现了解有限,因此系统评估和分析至关重要。本文通过扩展TEmBed对表格级嵌入进行系统评估,它能捕捉下游有效性所需的多种互补属性。对TEmBed模型池的实证研究证实,没有单一模型在所有任务中都表现出色,表明表格级嵌入质量不能仅归结为检索。
英文摘要
Tabular data is the dominant structured-data modality, and learning table representations has become a core research direction. Table-level embeddings in particular underpin a wide range of applications, including table retrieval, data lake discovery, and table classification. Despite their importance, there is still limited understanding of how different embedding approaches behave across tasks, making systematic evaluation and analysis essential. In this work, we introduce a systematic evaluation of table-level embeddings that captures several complementary properties required for downstream effectiveness. We realize this evaluation by extending TEmBed, a recently proposed testbed for tabular embeddings, whose table-level coverage is currently limited to a single retrieval task. An empirical study over the TEmBed model pool confirms that no single model excels across all tasks, demonstrating that table-level embedding quality cannot be reduced to retrieval alone.
CommentsAccepted to the 4th International Workshop on Tabular Data Analysis (TaDA) @ VLDB 2026