arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.24218cs.AIcs.CL

基于大语言模型的约束引导式企业数据映射

Constraint-Guided Enterprise Data Mapping with Large Language Models

Sebastian Monka, Pramod Anantharam, Thien Vo Minh, Lavdim Halilaj

首次发表
浏览论文内容

中文总结 AI 辅助

针对企业数据映射的问题,提出约束引导式映射(CGM)神经符号方法,通过三阶段流程提升匹配性能,该方法成本低、可迁移且能降低专家工作量,在基准测试中表现优异。

中文摘要 AI 辅助

企业实体匹配必须处理半结构化记录、隐式属性以及单位或粒度不匹配问题。实践中仍普遍采用人工匹配,但随着模式和数据源提供者的演变,人工匹配无法扩展。仅使用大语言模型(LLM)的匹配方法提升了语义召回率,但可能违反结构和物理不变量,生成流畅但操作上无效的对应关系。我们提出约束引导式映射(CGM),这是一种神经符号方法,包含三个阶段:(i)基于模式的可接纳性约束,元数据为mc = <τ_c, δ_c>,其中τ_c表示约束类型,δ_c提供可执行的关系和归一化逻辑;(ii)约束限制的候选生成,采用级联松弛以在存在噪声时保证非空可行集;(iii)在该可行集内进行有限LLM消歧的神经排序。从方法学角度看,约束作为假设空间算子而非事后验证器,使系统在松弛时可控制退化,并具备可审计、人类可引导的决策能力。在受控结构诱饵基准测试中,严格可接纳性将候选空间缩小约480倍且未丢失真实值(GT),逐层消融实验显示,该约束门而非LLM是性能提升的关键因素(F1从0.08升至0.66)。该方法的优势与模型无关且无额外推理成本:带约束的小模型以约28倍更低成本匹配了未使用约束的前沿LLM。该方法(非单一调优配置)可在7种企业数据源间迁移(宏F1为0.70),每种数据源都有其自动发现、专家可优化的约束,且相比电子表格工作流降低约7倍的专家工作量。公开的Valentine结果提供了外部排序合理性检查并划定了边界:仅在结构不变量决定匹配时,约束才应是严格的。

英文摘要

Enterprise entity alignment must handle semi-structured records, implicit attributes, and unit or granularity mismatches. Manual matching is still common in practice, but does not scale as schemas and providers evolve. LLM-only matching improves semantic recall, yet can violate structural and physical invariants, producing fluent yet operationally invalid correspondences. We propose constraint-guided mapping (CGM), a neuro-symbolic method with three stages: (i) schema-grounded admissibility constraints with metadata mc = <tau_c, delta_c>, where tau_c denotes the constraint type and delta_c provides executable relation and normalization logic; (ii) constraint-restricted candidate generation with cascade relaxation to guarantee a nonempty feasible set under noise; and (iii) neural ranking with bounded LLM disambiguation restricted to that feasible set. Methodologically, constraints operate as hypothesis-space operators rather than post-hoc validators, enabling controlled degradation under relaxation and auditable, human-guidable decisions. On a controlled structural-decoy benchmark, hard admissibility shrinks the candidate space by ~480x without dropping the GT, and a layer-by-layer ablation shows this gate, not the LLM, is the decisive lift (F1 0.08 to 0.66). The benefit is model-independent and adds no extra inference cost: a small model with constraints matches a frontier LLM used without them at ~28x lower cost. The method, not a single tuned configuration, transfers across seven enterprise makes (macro F1 0.70), each under its own automatically discovered, expert-refinable constraints, and lowers expert effort by ~7x versus spreadsheet workflows. Public Valentine results add an external ranking sanity check and mark the boundary: constraints should be hard only where structural invariants are match-determining.

发表机构

  • Bosch Center for Artificial Intelligence(博世人工智能中心)
  • Robert Bosch GmbH(罗伯特·博世有限公司)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑