发表机构
College of Mathematics and Physics, Suqian University(宿迁学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过两个对撞机工作流实验验证发现,潜空间物理可读性降低(泄漏减少)无法可靠预测轮廓似然偏差降低,表明潜可读性仅可作为路由诊断指标,不能作为推断鲁棒性的证明,需通过预设压力测试和留出集验证推断鲁棒性。
AI 中文摘要
可复用的对撞机表征可通过下游判别任务和保留信息探针进行评估,但这两个指标都无法直接检验轮廓似然中得分模板的行为。我们在受控的双通道路由协议中验证了一项具体预测:如果干扰分支中物理标签的可读性降低意味着表征的推断鲁棒性更强,那么在固定的未建模偏移下,它应伴随更小的轮廓信号强度偏差。\n在公开的Compact Muon Solenoid(紧凑μ子线圈)$H\rightarrow ZZ\rightarrow4\ell$ 工作流中,对固定EveNet嵌入进行下游拆分后,信号/背景的受试者工作特征曲线下面积($0.9894\pm0.0004$)得以保持,同时干扰分支的物理可读性从$0.961\pm0.013$降至$0.593\pm0.030$。探针灵敏度和有效秩对照排除了读出失败和分支崩溃的可能。在另一项顶夸克喷流标记工作流中,泄漏减少的现象再次出现且任务性能保持不变。\n然而,在两个开发事件分片上,它与最大绝对轮廓偏差的斯皮尔曼相关系数为$0.036$,且六次实质性泄漏改善的转换中有三次并未降低该偏差。在独立获取的分片上进行的一次性预注册验证显示,三组成对种子均实现了实质性泄漏减少,但其中两组的最大绝对偏差反而上升。因此,在测试的协议范围内,潜空间可读性是有用的路由诊断指标,但并非似然鲁棒性的证明。该结果支持一项实用验证规则:关于推断鲁棒性的主张需要预先指定的面向似然的压力测试和留出集验证。
英文摘要
Reusable collider representations can be evaluated through downstream discrimination and probes of retained information, but neither quantity directly tests the behaviour of score templates in a profiled likelihood. We test a specific prediction in a controlled two-channel routing protocol: if reduced physics-label readability in a nuisance branch indicates a more inference-robust representation, it should accompany a smaller profiled signal-strength bias under fixed unmodelled shifts. In a public Compact Muon Solenoid $H\rightarrow ZZ\rightarrow4\ell$ workflow, a downstream split of fixed EveNet embeddings preserves signal/background area under the receiver operating characteristic curve ($0.9894\pm0.0004$) while reducing nuisance-branch physics readability from $0.961\pm0.013$ to $0.593\pm0.030$. Probe-sensitivity and effective-rank controls exclude a failed readout and branch collapse. In a separate top quark jet-tagging workflow, the leakage reduction recurs with preserved task performance. Across two development event shards, however, its Spearman association with maximum absolute profiled bias is $0.036$, and three of six material leakage-improving transitions do not reduce that bias. A one-shot preregistered confirmation on an independently accessed shard produces material leakage reductions in all three paired seeds, while the maximum absolute bias increases in two. Thus, within the tested protocol, latent readability is a useful routing diagnostic but not a likelihood-robustness certificate. The result supports a practical validation rule: claims about inference robustness require a prespecified likelihood-facing stress test and held-out confirmation.
Comments27 pages, 6 figures