AI 中文总结
本文对6个分层程序语料库的内容近重复层、通用容器层及其联合进行对称审计,发现MyFixit和Doc2Dial呈单层通过/联合失败模式,完成了有界测量与审计,未提出新算法或因果主张。
AI 中文摘要
组件不相交泄漏控制可通过内容相似性、层级成员关系或两者对语料库单元进行分组。Guvenilir与Dogan此前已证明,合并关系类型会生成一个阻碍拆分的巨型组件,我们不主张该现象为新发现。我们通过对内容近重复层、通用容器层及其联合的对称审计,在固定的6个分层程序语料库面板中展开研究。MyFixit和Doc2Dial在相同操作标准下呈现出“单层通过/联合失败”模式,由此得到的2/6比例描述该刻意构建的面板,而非流行度估计。预先指定的特定桥接预测器与该模式相关,但未与已登记的联合密度控制区分开,因此该面板无法识别特定桥接机制。次要诊断约束了阈值敏感性、注释覆盖范围及词汇线索的解释,且未将这些发现扩展至其记录来源与定义之外。本文的贡献在于有界测量与审计:它保持关系族可见,对称评估其个体及联合组件结构,并报告负例与机制限制。本文未提出新的拆分算法,也未对关系联合提出一般性因果主张。
英文摘要
Component-disjoint leakage control can group corpus units by content similarity, hierarchical membership, or both. Guvenilir and Dogan previously showed that merging relation types can create a giant component that obstructs splitting; we do not claim this phenomenon as new. We examine it through a symmetric audit of a content-near-duplicate layer, a common-container layer, and their union in a fixed panel of six hierarchical procedural corpora. MyFixit and Doc2Dial exhibit the individual-layer-pass/union-fail pattern under the same operational criteria. The resulting two-of-six fraction describes this deliberately constructed panel and is not a prevalence estimate. A prespecified bridge-specific predictor is associated with the pattern, but it is not distinguished from a registered union-density control; the panel therefore does not identify a bridge-specific mechanism. Secondary diagnostics bound the interpretation of threshold sensitivity, annotation coverage, and lexical cues without extending those findings beyond their recorded sources and definitions. The paper's contribution is a bounded measurement and audit: it keeps relation families visible, evaluates their individual and union component structures symmetrically, and reports negative cases and mechanism limits. It proposes neither a new splitting algorithm nor a general causal claim about relation unions.
Comments19 pages, 1 figure, 9 tables