AI 中文总结
研究人工智能证据开源项目中架构变更后稳定义务衡量问题,指出后锚定稳定密度指标失败,引入稳定机制并人工验证校准,强调需受控设计与机制感知归因来可靠识别稳定义务。
AI 中文摘要
仓库挖掘研究越来越多地分析人工智能证据项目,但尚不清楚如何衡量架构变更是否产生了延迟的稳定义务。一种自然指标,即后锚定稳定密度,计算在引入持久边界后出现的测试、持续集成门、文档和修复。我们表明该指标失败。在对338个2026年高知名度GitHub仓库的差异级别研究中,锚定很常见,但一个受控的308事件连续窗口实验发现没有后锚定提升。我们引入了稳定机制,即六种反复出现的模式,解释了密度指标为何失败,并通过人工验证对其进行校准。验证揭示了一个两层陷阱。后锚定密度将普遍维护与真正的债务混为一谈;在仓库挖掘能够可靠地识别稳定义务之前,具有机制感知归因的受控设计是必要的。
英文摘要
Repository mining studies increasingly analyze AI-evidence projects, yet it remains unclear how to measure whether architectural changes create deferred stabilization obligations. A natural metric, post-anchor stabilization density, counts tests, CI gates, documentation, and fixes appearing after a durable boundary is introduced. We show that this metric fails. In a diff-level study of 338 high-visibility 2026 GitHub repositories, anchors are common (321 of 338 contain real changed-file anchor evidence), but a controlled 308-event contiguous-window experiment finds no post-anchor uplift: broad and strict stabilization signals both yield median post/pre density ratios near 1.0, and non-anchor controls are equally dense. We introduce stabilization regimes, six recurring patterns that explain why the density metric fails, and use human validation to calibrate them. Two independent coders label 100 stratified candidate-anchor events from blinded packets (kappa = 0.50 on debt attribution). The validation exposes a two-layer trap: many candidate anchors are not durable boundaries (45 of 100), and even among valid anchors in this calibration sample the no-uplift result holds: only 3 of 100 events survive as attributable delayed obligations; the remaining 97 are explained by classifier error, anchor-local hardening, background maintenance, or pre-anchor hardening. Post-anchor density conflates pervasive maintenance with genuine debt; controlled designs with regime-aware attribution are necessary before repository mining can reliably identify stabilization obligations.