AI 中文总结
该研究提出CASE框架,结合LLMs语义理解与表格学习器统计能力,通过预填充Gemma 3表格语言模型KV缓存建立语义锚点,在CARTE等基准上显著提升表格学习器在语义丰富数据集的低数据场景性能。
AI 中文摘要
尽管现代表格学习器擅长捕捉统计模式,但它们常处于语义真空状态,将文本特征视为离散符号,忽略特征名称或单元格条目中蕴含的丰富语义。我们提出CASE(Context-Aware Semantic Embeddings,上下文感知语义嵌入),这是一种弥合大型语言模型(LLMs)语义理解能力与表格学习器统计能力之间差距的新型框架。与现有孤立嵌入行的方法不同,CASE采用上下文策略:我们用代表性行样本预填充自定义训练的基于Gemma 3的表格语言模型的KV缓存,以建立数据集语义的持久锚点。这确保生成的行嵌入是动态上下文化的,解决语义歧义并将表示锚定在特定领域上下文中。我们在多个基准(CARTE、TextTab和TabArena)上的实验表明,CASE可显著提升表格学习器在语义丰富数据集上的性能,尤其在低数据 regime( regime 译为“ regime”,此处保留原词)下表现突出。
英文摘要
While modern tabular learners excel at capturing statistical patterns, they frequently operate in a semantic vacuum, treating textual features as discrete symbols, ignoring the rich semantics inherent in feature names or cell entries. We propose CASE (Context-Aware Semantic Embeddings), a novel framework that bridges the gap between the semantic understanding of Large Language Models (LLMs) and the statistical capabilities of tabular learners. Unlike existing methods that embed rows in isolation, CASE utilizes a contextualization strategy: we pre-fill the KV cache of a custom-trained Gemma 3-based Tabular Language Model with a representative sample of rows to establish a persistent anchor of the dataset's semantics. This ensures that generated row embeddings are dynamically contextualized, resolving semantic ambiguities and anchoring representations in domain-specific context. Our experiments across several benchmarks (CARTE, TextTab, and TabArena) demonstrate that CASE substantially improves the performance of tabular learners on semantically rich datasets, particularly in low-data regimes.