arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.15205cs.DBcs.AI

TEAR:通过大语言模型从文本中进行带属性推荐的表格提取

TEAR: Table Extraction with Attribute Recommendation from Texts via Large Language Models

Tong Li, Shuye Ding, Jiachuan Wang, Yongqi Zhang, Shuangyin Li, Lei Chen, Bo Li

首次发表
浏览论文内容

中文总结 AI 辅助

TEAR框架通过动态调整指令和属性推荐工作流,解决自然文本中表格提取的边界划定和未知属性发现难题,实现最先进性能。

中文摘要 AI 辅助

从文本中提取表格是信息系统的一项重要任务,近年来,通过指令提示大语言模型(LLMs)的方法因其强大的性能而受到广泛关注。现有工作假设输入文本是表格描述或专门文档。然而,这些工作在很大程度上忽视了另一类常见于新闻报道和社交媒体中的文本类别:自然发生的文本。从这类文本中提取表格信息面临两个不同的挑战。首先,高变异性和缺乏显式结构线索使得固定的启发式LLM提示在精确划定提取边界方面能力有限。其次,手动预定义的模式无法捕捉自然发生文本中开放式的、未见过的属性。在本文中,我们提出了一个框架TEAR来应对这些挑战。它包含两个协同工作流:一个表格提取工作流,动态调整指令以克服启发式指令的局限性;以及一个属性推荐工作流,从文本中发现新属性以补充启发式模式。据我们所知,TEAR是第一个支持自动化文本驱动属性推荐的框架,能够实现表格提取的探索性模式设计。为了评估TEAR,我们建立了针对自然发生文本的表格提取和属性推荐基准,包括两个真实世界数据集、人工标注、适当的评估指标和基线比较。实验表明,TEAR在这两项任务上均达到了最先进的性能,并且推荐的属性在探索性场景中有效提升了提取性能。

英文摘要

Table extraction from texts is an important task for information systems, and recent approaches that prompt large language models (LLMs) with instructions have drawn great attention for their strong performance. Existing works have assumed the input texts to be table descriptions or specialized documents. However, these efforts have largely overlooked another prevalent category of texts, commonly found in news reports and social media: naturally occurring texts. Extracting tabular information from such texts poses two distinct challenges. First, high variability and the absence of explicit structural cues make fixed heuristic LLM prompts limited in precisely delineating extraction boundaries. Second, manually predefined schemas cannot capture open-ended, unseen attributes in naturally occurring text. In this paper, we propose a framework, TEAR, to address these challenges. It comprises two synergistic workflows: a Table Extraction Workflow that dynamically adapts instructions to overcome the limitation of heuristic instructions, and an Attribute Recommendation Workflow that discovers new attributes from texts to complement the heuristic schema. To our knowledge, TEAR is the first framework that supports automated text-driven attribute recommendation, enabling exploratory schema design for table extraction. To evaluate TEAR, we establish the benchmark for table extraction and attribute recommendation on naturally occurring texts, including two real-world datasets, manual annotations, appropriate metrics, and baseline comparisons. Experiments show that TEAR achieves state-of-the-art performance on both tasks, and the recommended attributes effectively enhance extraction performance in exploratory scenarios.

发表机构

  • Hong Kong University of Science and Technology(香港科技大学)
  • Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
  • South China Normal University(华南师范大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑