arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.07946cs.DBcs.AI

仅从值进行元数据重构:在无文档数据仓库中恢复列语义

Metadata Reconstruction from Values Alone: Recovering Column Semantics in Undocumented Warehouses

  • Independent Researcher USA
  • Independent Researcher

机构由 AI 辅助整理,请以论文原文为准。

Mike Helwig

AI总结:

本研究提出Rosetta框架,结合确定性分析器与语言模型,在无文档数据仓库中仅从值恢复列语义,经BIRD、i2b2等数据集验证,提升了元数据恢复的准确率与覆盖率。

AI中文摘要:

Text-to-SQL基准测试提供的模式中,列名已表明列的含义,而生产数据仓库则相反:存在模糊标识符、部分或缺失文档。我们首先解决由此产生的问题:仅从数据本身恢复列和值的含义。Rosetta将语言模型置于验证框架内:确定性分析器提取结构证据(值指纹、26种模式库、校验和判定),模型基于该证据提出语义,且每个事实都带有来源信息和受证据类别约束的置信度。在11个BIRD数据库的680对列(标识符已被破坏)上,针对人工文档,该框架在其确定的42%列上的元数据准确率为0.475,而同一模型直接使用时在94%列上的准确率为0.223。在两者均能处理的283列中,该框架生成的文本与单独模型相比并无更优;增益在于选择:确定性证据决定系统是否输出(覆盖率提升0.257[0.128,0.378]),而非输出质量。确定性层是能力检测器,而非能力放大器。骨干替换限定了该结论:文本发现可复现,但提示请求的弃权(不执行)无法迁移;代码强制提交门(先注册预测,在第三个骨干及保留数据库上测量)使所有骨干的无证据覆盖率为0.000。在盲测的i2b2临床数据仓库中,Rosetta仅从值就解码了134个真实ICD-9代码的95.5%,并对所有44个NDC药物代码弃权(不执行)。该目录支持查询时的校准弃权(不执行):在完全模式不透明情况下,朴素翻译器的执行准确率从0.92降至0.42,而我们的门在59%的覆盖率上达到86%的准确率。负面结果被明确报告,包括我们自身的权威阶梯并非该标题背后的机制。

英文摘要:

Text-to-SQL benchmarks ship schemas whose column names already say what the columns mean. Production warehouses are the inverse: cryptic identifiers, partial or absent documentation. We address the problem they pose first: recovering what columns and values mean from the data itself. Rosetta places a language model inside a verification harness: a deterministic profiler extracts structural evidence (value fingerprints, a 26-pattern library, checksum verdicts), the model proposes semantics conditioned on that evidence, and every fact carries provenance and a confidence bounded by its evidence class. Against human documentation on 680 paired columns across eleven BIRD databases, identifiers destroyed, the harness delivers metadata that is 0.475 accurate on the 42% of columns it commits to, against 0.223 on 94% for the same model used directly. Restricted to the 283 columns where both arms speak, the harness writes no better prose than the model alone; the gain is selection: deterministic evidence governs whether the system speaks (coverage +0.257 [0.128, 0.378]), not how well. The deterministic layer is a competence detector, not a competence amplifier. A backbone swap bounds the claim: the prose finding reproduces, but prompt-requested abstention does not transfer; a code-enforced commit gate (predictions registered first; measured on a third backbone and held-out databases) makes no-evidence coverage 0.000 on every backbone. On a blind i2b2 clinical warehouse Rosetta decodes 95.5% of 134 real ICD-9 codes from values alone and abstains on all 44 NDC drug codes. The catalog supports calibrated abstention at query time: under full schema opacity a naive translator falls from 0.92 to 0.42 execution accuracy while our gate answers at 86% accuracy over 59% coverage. Negative results are reported plainly, including that our own authority ladder is not the mechanism behind the headline.

↑