开放表格洞察提取:我们处于何种位置,又应走向何方?
Open Tabular Insight Extraction: Where Do We Stand, and Where Should We Go?
- Centrum Wiskunde & Informatica(数学与计算机科学研究中心)
- University of Amsterdam(阿姆斯特丹大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出开放表格洞察提取(OpenTI)框架,整合多社区术语,系统回顾现有系统与基准,指出其未覆盖端到端范围且不适于开放评估,并提炼研究议程以推动该领域发展。
AI中文摘要:
民主化获取数据湖等大型表格语料库中所蕴含的知识,正成为一个核心研究挑战。该领域的研究正在推进并拓宽范围,日益提供端到端满足个人洞察需求的组件。然而,这些努力在不同社区之间仍然分散,各社区以自身惯例来界定问题,例如表格问答、文本转SQL和数据分析代理,其中作品引用同一任务标签内文献的可能性是跨标签引用的六倍。为使这些社区达成共识,我们为这一追求建立了一个整体框架,称之为开放表格洞察提取(OpenTI)。我们从第一性原理出发,围绕个人所需的分析知识、从表格语料库中推导该知识的程序,以及结果对寻求者服务的满意程度,对OpenTI进行形式化。在此过程中,我们整合了信息检索、自然语言处理、机器学习、数据库和人机交互领域的框架与术语,并将这一基础应用于对致力于OpenTI的系统与基准的系统性回顾与分析。我们发现,当前系统并未覆盖OpenTI的端到端范围,主要聚焦于分析本身,而基准在很大程度上不适合开放环境下的评估,因为输入预设了对表格的知识,且验证机制与设置不匹配。最后,我们提炼出一项研究议程,旨在推动OpenTI系统、评估及交互范式的发展,以揭示用户所需的洞察。我们论文的交互式配套材料可在此https URL获取。
英文摘要:
Democratizing access to the knowledge held in large corpora of tables such as data lakes is emerging as a central research challenge. Research in this space is advancing and broadening in scope, increasingly supplying the components to satisfy a person's insight need end-to-end. Yet these efforts remain fragmented across communities that frame the problem under their own conventions, such as table question answering, text-to-SQL, and data analysis agents, with works six times as likely to cite within the same task label as across labels. To bring these communities onto common ground, we establish a holistic framework for this pursuit, which we refer to as Open Tabular Insight Extraction (OpenTI). We formalize OpenTI from first principles around the analytical knowledge a person needs, the procedure for deriving it from a corpus of tables, and how well a result serves the person who sought it. In doing so we consolidate frameworks and terminology across information retrieval, natural language processing, machine learning, databases, and human-computer interaction, and apply this grounding in a systematic review and analysis of systems and benchmarks that work towards OpenTI. We find that current systems do not cover the end-to-end scope of OpenTI, mainly focusing on the analysis itself, and that benchmarks are largely unfit for evaluations in an open setting as inputs presuppose knowledge of tables, and validation mechanisms do not match the setup. Finally, we distill a research agenda towards OpenTI systems, evaluation, and interaction paradigms that surface the insights users need. An interactive companion to our paper is available at https://open-tabular-insight-extraction.github.io.