arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.04269cs.DBcs.IRcs.LGcs.LO

企业家族关系解析并非字符串匹配问题:按名称可见性分层的公共基准

Corporate-Family Resolution Is Not a String-Matching Problem: A Public Benchmark Stratified by Name Visibility

Harshit Gupta

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出按名称可见性分层的公共基准CorpFam,发现企业家族关系解析是检索问题而非匹配问题,干预点为候选生成,不可见名称对的匹配性能极低,相关基准与代码已公开。

中文摘要 AI 辅助

判断两条供应商记录是否属于同一企业家族是支出整合、信用风险汇总和制裁筛查的前提,该任务通常被当作实体匹配处理,但二者存在差异:家族关系连接的是刻意不同的实体,证据往往未出现在任何一条记录中。我们推出CorpFam,这是一个公共基准,包含来自6638350条美国联邦授标记录的10307个企业家族的54864个候选对,其中每个供应商都向政府登记处自报其最终母公司。这些对按名称可见性分层:归一化后名称相同、共享独特标记或无共享标记。由于各层的正例率在10.2%至97.3%之间,我们报告分层召回率(与基准率无关),而非F1值(与基准率相关)。5种匹配器中最强的一种可恢复100.0%的相同名称对和4.2%的不可见名称对;没有方法在后者上超过4.7%。失败在匹配阶段之前就已出现:分块决定匹配器可见的对,我们评估了7种方案,包括语音键、忽略名称的属性键和语义最近邻。没有一种方案在不可见名称对上达到3%,它们的并集可恢复6.8%。93.2%的此类链接从未进入候选集,因此匹配阶段的改进无法触及它们。这些链接是真实的:与来源不同的SEC附表21子公司表相比,64.2%的不可见链接得到证实,而置换母公司的情况下为0.16%,同一母公司的错误附表情况下为0.41%,两个无关空值的一致性在0.25分以内。企业家族关系解析是被误归类为匹配问题的检索问题;干预点是候选生成,而非排序。该基准、裁决日志和复现所有数值的代码已发布。

英文摘要

Deciding whether two supplier records belong to the same corporate family is a prerequisite for spend consolidation, credit exposure aggregation and sanctions screening. It is usually treated as entity matching, but the tasks differ: a family link connects records that are deliberately different entities, and the evidence often appears in neither record. We introduce CorpFam, a public benchmark of 54,864 candidate pairs over 10,307 corporate families, derived from 6,638,350 US federal award records in which every supplier self-reports its ultimate parent to a government registry. Pairs are stratified by name visibility: whether the names are identical after normalisation, share a distinctive token, or share none. Because strata have positive rates from 10.2% to 97.3%, we report per-stratum recall, base-rate invariant, rather than F1, which is not. The strongest of 5 matchers recovers 100.0% of identical pairs and 4.2% of invisible ones; no method exceeds 4.7% on the latter. The failure begins before matching. Blocking decides which pairs a matcher sees, and we evaluate 7 schemes spanning phonetic keys, attribute keys that ignore the name, and semantic nearest neighbours. None reaches three percent on invisible pairs, and their union recovers 6.8%. 93.2% of these links never enter the candidate set, so no matching-stage improvement can reach them. The links are real: against SEC Exhibit 21 subsidiary schedules, which share no provenance with procurement registration, 64.2% of invisible links are corroborated, against 0.16% under permuted parents and 0.41% against the same parent's wrong exhibit: two unrelated nulls agreeing to within 0.25 points. Corporate-family resolution is a retrieval problem misfiled as a matching problem; the intervention point is candidate generation, not ranking. The benchmark, adjudication log, and code reproducing every number are released.

补充信息

↑