Glyph:用于企业数据目录列描述与敏感性本体标注的多策略智能体系统
Glyph: A Multi-Strategy Agentic System for Column Description and Sensitivity-Ontology Tagging of Enterprise Data Catalogs
浏览论文内容
中文总结 AI 辅助
Glyph是一个多策略智能体系统,通过描述器与标注器协作,利用主动检索增强生成和三种并行策略融合,自动生成列描述并标注敏感性本体,显著提升标注质量与可审计性。
中文摘要 AI 辅助
企业数据湖中表的积累速度远超人工管理员记录或分类的速度,导致列缺少描述且未分配治理标签。这种文档负债损害了数据发现、访问控制和法规遵从性。我们提出了Glyph,一个生产系统,将两个耦合问题——列描述生成和用于数据分类的列类型标注——构建为协作的LLM智能体,以有状态图的形式编排。描述器(Descriptor)将生成过程锚定在产生每列的流水线源代码中,通过推理-行动工具循环(主动检索增强生成)按需从企业GitHub检索。标注器(Tagger)通过并行运行三种互补策略(描述标注器、业务线正则表达式标注器和由微调对比编码器支持的元数据标注器,基于向量数据库),从受治理的275叶数据分类本体中分配标签,然后使用倒数排名融合(RRF)融合其排名输出。我们使用批内对比目标微调了一个6层MiniLM元数据编码器,将分布内留出集上的同标签检索从NDCG@10 0.55提升到0.92(MAP@100从0.19提升到0.90),相对于基础编码器。我们报告了在三个评估组中,在召回加权F2目标下的端到端多标签标注质量,一个隔离每种策略和RRF融合的消融实验,以及将Glyph与先前列类型标注工作和商业价值/正则表达式敏感性扫描器区分开来的工程决策:无值且基于代码的设计、每个标签的来源追踪以及优雅降级。这些共同使多智能体LLM目录编目可审计,并可作为生产服务运营。
英文摘要
Enterprise data lakes accumulate tables faster than human stewards can document or classify them, leaving columns with missing descriptions and unassigned governance labels. This documentation debt undermines data discovery, access control, and regulatory compliance. We present Glyph, a production system that frames two coupled problems, column description generation and column type annotation for data classification, as cooperating LLM agents orchestrated as stateful graphs. The Descriptor grounds generation in the pipeline source code that produces each column, retrieved on demand from an enterprise GitHub via a reasoning--acting tool loop (active Retrieval-Augmented Generation). The Tagger assigns labels from a governed 275-leaf Data Classification Ontology by running three complementary strategies in parallel (a description tagger, a line-of-business regex tagger, and a metadata tagger backed by a fine-tuned contrastive encoder over a vector database), then fuses their ranked outputs with Reciprocal Rank Fusion (RRF). We fine-tune a 6-layer MiniLM metadata encoder with an in-batch contrastive objective, lifting same-tag retrieval on an in-distribution held-out split from NDCG@10 0.55 to 0.92 (MAP@100 $0.19 \rightarrow 0.90$) relative to the stock base encoder. We report end-to-end multi-label tagging quality under a recall-weighted F2 objective across three evaluation groups, an ablation isolating each strategy and the RRF fusion, and the engineering decisions that distinguish Glyph from prior column-type-annotation work and from commercial value/regex sensitivity scanners: value-free and code-grounded design, per-tag provenance, and graceful degradation. Together these make multi-agent LLM cataloging auditable and operable as a production service.