arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.23767cs.DLcs.DB

读回数据:丰富变量级元数据以进行模型-数据一致性检查

Reading the Data Back: Enriching Variable-Level Metadata for Model-Data Consistency Checks

Eryk Kulikowski

首次发表
浏览论文内容

中文总结 AI 辅助

针对研究数据存储库缺乏变量级元数据的问题,提出兼容DDI-CDI的元数据配置文件,实现模型-数据一致性检查,并通过筛选4,868个复制数据集验证其有效性。

中文摘要 AI 辅助

目的:研究数据存储库通常缺乏变量级元数据。模型选择筛选,即检查所报告的模型是否适合其结果变量的取值,也需要这些值的摘要以及使用这些值的分析链接。我们探讨存储库如何以带有证据和审查历史的方式表示和获取这些元数据。方法:我们提出一个与数据文档倡议跨域集成(DDI-CDI)模型兼容的元数据应用配置文件,将版本化的变量和经验概况链接到所报告的分析、估计器及分析角色。提取来源和审查决定被分别记录。我们评估了流式数据概况分析、从Stata和R代码中确定性提取,以及从论文中提取分析记录的语言模型。结果:一个复制包的工作导出示例说明了该配置文件:它通过了形状和交换检查,并回答了四个查询,其中一个返回未引发警报的分析。对来自六种政治学期刊的4,868个复制数据集进行筛选,将结果链接到1,440个存储中的概况列,并标记了965个,包括21个具有手工验证的计数、比例、二元或有序结果的线性模型存储中的14个。AI辅助裁决在119个计数和比例候选者的近平衡样本上产生了55%的精确度;筛选规则是在同一语料库上开发的。结论:该配置文件使分析-变量关系可查询,同时保留证据和审查历史;筛选产生带有上下文的审查候选。代码、元数据和测量结果已发布。

英文摘要

Purpose: Research data repositories often lack variable-level metadata. Model-choice screening, checking whether a reported model suits the values of its outcome variable, also requires summaries of those values and links to the analyses that use them. We ask how repositories can represent and acquire them as metadata with evidence and review histories. Methods: We propose a metadata application profile compatible with the Data Documentation Initiative Cross-Domain Integration (DDI-CDI) model, linking versioned variables and empirical profiles to reported analyses, estimators, and analysis roles. Extraction provenance and review decisions are recorded separately. We evaluate streaming data profiling, deterministic extraction from Stata and R code, and language-model extraction of analysis records from papers. Results: A worked export of one replication package illustrates the profile: it passes shape and interchange checks and answers four queries, one of which returns the analyses that raise no alert. Screening 4,868 replication datasets from six political-science journals links outcomes to profiled columns in 1,440 deposits and flags 965, including 14 of the 21 deposits with a hand-verified linear model on a count, proportion, binary, or ordinal outcome. AI-assisted adjudication yields 55% precision on a near-balanced sample of 119 count and proportion candidates; screening rules were developed on the same corpus. Conclusion: The profile makes analysis-variable relationships queryable while preserving evidence and review history; screening yields review candidates with context. Code, metadata, and measurements are released.

发表机构

  • KU Leuven(荷语鲁汶大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑