arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.20887cs.LGcs.AImath.AG

TwistedMerge:用于模型合并的经过认证的高阶诊断与弃权

TwistedMerge: Certified Higher-Order Diagnostics and Abstention for Model Merging

  • Department of Mathematics, University of Washington(华盛顿大学数学系)
  • Beijing International Center for Mathematical Research, Peking University(北京大学北京国际数学研究中心)

机构由 AI 辅助整理,请以论文原文为准。

Ting Gong, Shitan Xu

中文总结 AI 辅助

研究模型合并中全局对齐问题,提出TwistedMerge认证管道,将合并视为有限下降问题,经多测试提升残差为上同调类,证明相关定理,通过实验展示该方法在处理多种情况时的表现,定位下降理论为可证伪框架。

中文摘要 AI 辅助

模型合并会结合独立训练或微调的模型,但成对可对齐性并不意味着全局一致对齐。我们将合并表述为一个有限下降问题,其中检查点是局部对象,对齐映射是转换,循环积是残差。TwistedMerge是一个保守的认证管道,它分离了固定图平均、可同步消除的规范不一致、在指定比较复形上的经过认证的中心障碍以及非阿贝尔全纯性。只有在经过反一致性、系数识别、中心性和闭包测试后,残差才会被提升为上同调类;否则该方法弃权并返回普通或同步回退。我们证明了常数边不可行结果、冻结复形三元组和预声明族错误控制定理以及比较复形灵敏度的细化测试。通过循环一致同步消除了植入的神经对齐缺陷,表明仅非零循环分数不是更高的障碍。受控中心系统恢复了预测的非上边缘和射影秩行为,而噪声估计在测试控制上从认证转变为弃权且无错误提升。一个经过训练的低秩适配器审计表明,朴素因子平均取决于所选的GLr代表,而全局因子同步和密集增量奇异值分解是稳定的。在自然检查点集合上,循环残差不能预测合并退化,并且没有自然的中心或周期索引类被认证。这些结果将下降理论定位为一个可证伪的认证和弃权框架。

英文摘要

Model merging combines independently trained or fine-tuned models, but pairwise alignability does not imply globally consistent alignment. We formulate merging as a finite descent problem in which checkpoints are local objects, alignment maps are transitions, and cycle products are residuals. TwistedMerge is a conservative certification pipeline that separates fixed-chart averaging, synchronization-removable gauge inconsistency, a certified central obstruction on a specified comparison complex, and nonabelian holonomy. A residual is promoted to a cohomology class only after inverse-consistency, coefficient-identification, centrality, and closure tests; otherwise the method abstains and returns an ordinary or synchronized fallback. We prove a constant-edge no-go result, frozen-complex three-way and predeclared-family error-control theorems, and a refinement test for comparison-complex sensitivity. A planted neural alignment defect is removed by cycle-consistent synchronization, showing that a nonzero cycle score alone is not a higher obstruction. Controlled central systems recover the predicted non-coboundary and projective-rank behavior, while noisy estimates move from certification to abstention without false lifts on the tested controls. A trained low-rank-adapter audit shows that naive factor averaging depends on the chosen GLr representative, whereas global factor synchronization and dense-delta SVD are stable. On natural checkpoint collections, cycle residuals do not predict merge degradation and no natural central or period-index class is certified. The results position descent theory as a falsifiable certification and abstention framework.

补充信息

↑