arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从研究问题到列:感知操作化的数据发现

From Research Questions to Columns: Operationalization-Aware Data Discovery

Houming Chen, H. V. Jagadish

arXiv 2608.04536首次发表:更新:

AI 中文总结

该研究提出感知操作化的数据发现(OADD)任务,构建基准OADD-Bench,评估多种方法后发现OADD仍是未解决问题,其中OADD定向智能体表现最佳。

AI 中文摘要

研究人员通常带着一个抽象概念访问数据仓库,询问哪些列可以衡量该概念。有用的列可能与查询不相似,它们可能仅作为可辩护衡量标准中的补充指标才重要。这种需求不同于模式链接和列检索,后两者从更明确的需求出发,奖励直接相关性。我们定义了感知操作化的数据发现(OADD):给定一个宽泛问题和一个数据库(可选受范围约束),OADD共同确定如何用可用数据衡量核心概念,并识别支持性列。开发OADD方法需要用于设计和评估的示例,但要求研究人员提供概念问题及其对应列是不切实际的。我们通过将实证论文视为使用中的模式记录来构建OADD-Bench。问题挖掘器提取并重构论文支持的问题;论文条件列挖掘器重建其衡量标准并将其基于数据库标识符。我们仅接受出版物和数据库文档支持的映射。OADD-Bench包含来自111篇论文的160个问题和4682个问题-列标签。每个目标记录已发表研究中使用的衡量标准;论文提供先例,而挖掘器提取并将其基于数据库标识符。我们评估了词汇检索、神经检索、适配的模式链接系统和大语言模型(LLM)OADD智能体。每种方法仅接收一个问题、允许的年份和数据集元数据;源论文仅用于构建和记录基准标签。在最大输出限制下,直接检索的召回率最多为0.185。最强的模式链接适配达到0.401,但仍针对不同目标优化;OADD定向智能体表现最佳,达到0.465。即使该智能体覆盖的目标列不到一半,这表明OADD仍然是一个未解决的问题。

英文摘要

Researchers often approach a data repository with an abstract concept and ask which columns can measure it. Useful columns may not resemble the query; they may matter only as complementary indicators in a defensible measure. This need differs from schema linking and column retrieval, which begin from more explicit needs and reward direct relevance. We define operationalization-aware data discovery (OADD): given a broad question and a database, optionally under a scope constraint, OADD jointly determines how focal concepts can be measured with available data and identifies supporting columns. Developing OADD methods requires examples for design and evaluation, but asking researchers to supply conceptual questions and their columns is impractical. We construct OADD-Bench by treating empirical papers as records of schema in use. A question miner extracts and reframes a paper-supported question; a paper-conditioned column miner reconstructs its measurements and grounds them to database identifiers. We admit only mappings supported by the publication and database documentation. OADD-Bench contains 160 questions from 111 papers and 4,682 question-column labels. Each target records a measurement used in published research; the paper supplies the precedent, while the miners extract and ground it. We evaluate lexical and neural retrieval, adapted schema-linking systems, and large language model (LLM) OADD agents. Each method receives only a question, permitted years, and dataset metadata; source papers are used only to construct and document benchmark labels. At the largest output limit, direct retrieval reaches at most 0.185 recall. The strongest schema-linking adaptation reaches 0.401 but remains optimized for a different objective; an OADD-directed agent performs best at 0.465. Even this agent covers less than half the target columns, showing that OADD remains an open problem.

Comments6 pages, 2 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑