发表机构
Harbin Institute of Technology(哈尔滨工业大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过大规模实验证明,近似函数依赖强度度量与错误检测价值(精确率/召回率)存在显著非对称性,强度排序不可靠,二者应分别评估。
AI 中文摘要
近似函数依赖(AFD)强度度量通常用于对候选依赖进行优先级排序,包括当所选依赖支持错误检测等下游任务时。这种做法隐含地假设,按强度排名较高的依赖对于其旨在支持的任务而言也是更好的候选。我们直接针对错误检测检验了这一假设。我们通过精确率和召回率来评估检测价值,而不对不同类型的错误施加单一权重,并且仅在需要单一候选选择目标时引入显式的假阳性/假阴性成本。在来自一个成熟AFD基准的三个真实世界关系、1,155次损坏与检测运行以及513,480次候选级评估中,我们发现了一个明显的非对称性。强度度量倾向于与精确率呈正相关,但它们与召回率的关联则远不稳定,其幅度和方向在不同度量和实验设置中均有变化。基于强度的优先级排序在排名层面也不可靠:按强度选择的top-k集合与按实际检测性能所青睐的集合往往重叠很少,且基于强度的选择可能产生显著的遗憾。召回率不匹配的一个重要部分与候选等价类划分下的结构检测机会相关,后者在整体上比我们评估的强度度量更一致地追踪召回率。损坏机制和若干静态结构属性可以解释选定的效应,但无法提供跨设置的共同解释,留下了部分不匹配问题未解决。总体而言,AFD强度与检测价值是不同的,应分别进行评估。
英文摘要
Approximate functional dependency (AFD) strength measures are commonly used to prioritize candidate dependencies, including when selected dependencies support downstream tasks such as error detection. This practice implicitly assumes that dependencies ranked higher by strength are also better candidates for the task they are intended to support. We test this assumption directly for error detection. We evaluate detection value through precision and recall without imposing a single weighting on different error types, and introduce an explicit false-positive/false-negative cost only when a single candidate-selection objective is required. Across three real-world relations from an established AFD benchmark, $1{,}155$ corruption-and-detection runs, and $513{,}480$ candidate-level evaluations, we find a clear asymmetry. Strength measures tend to track precision positively, but their associations with recall are substantially less stable, with both magnitude and direction varying across measures and experimental settings. Strength-based prioritization is also unreliable at the ranking level: top-$k$ sets chosen by strength often have little overlap with those favored by actual detection performance, and strength-based selections can incur substantial regret. An important part of the recall mismatch is associated with structural detection opportunity under the candidate equivalence-class partition, which tracks recall more consistently overall than the strength measures we evaluate. Corruption mechanism and several static structural properties account for selected effects but do not provide a common explanation across settings, leaving part of the mismatch unresolved. Overall, AFD strength and detection value are distinct and should be evaluated separately.