发表机构
Korea University(韩国延世大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
PEARL是无需训练的多表格检索框架,通过前置化关系链实现垂直划分的子表格编码,在3跳查询的R@2指标上最高提升30.05%,性能优于现有方法。
AI 中文摘要
尽管大语言模型(LLMs)已展现出强大的表格推理能力,但由于现实世界数据的碎片化和关系型结构,检索相关表格仍具挑战性。现有工作通常依赖全表表示,忽略了由连接关系(join relationships)产生的跨表格语义。我们提出PEARL,这是一个无需训练的框架,将范式转向基于垂直划分的子表格编码。PEARL通过在预先识别的连接路径上生成多跳查询,离线扩充检索语料库,并将相关列重组为垂直划分的语料库单元,实现无需查询时LLM推理的有效多表格检索。实验表明,PEARL在3跳查询上的R@2指标相比现有方法始终表现更优,提升幅度最高达30.05%。源代码可通过此URL获取。
英文摘要
While large language models (LLMs) have shown strong capabilities in tabular reasoning, retrieving relevant tables remains challenging due to the fragmented and relational structure of real-world data. Existing work typically relies on whole table representations that overlook cross-table semantics induced by join relationships. We propose PEARL, a training-free framework that shifts the paradigm toward vertical partitioning-based sub-table encoding. PEARL augments the retrieval corpus offline by generating multi-hop queries over pre-identified join paths and reorganizing relevant columns into vertically partitioned corpus units, enabling effective multi-table retrieval without query-time LLM inference. Experiments show that PEARL consistently outperforms existing methods, with up to +30.05% gains in R@2 on 3-hop queries. The source code is available at https://github.com/SOOB2NHO/PEARL.
CommentsAccept to EMNLP 2026