发表机构
PolymathMinds Lab; aSSIST University; Samsung Engineering(多智思维实验室; aSSIST大学; 三星工程公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对专家领域树形结构与欧几里得变换器的问题,提出HySAT方法,即仅在损失层使用双曲损失。通过构建和部署六个专家语言模型进行实验,证明该方法的有效性及稳定性,四个模型已实际应用,有相关记录。
AI 中文摘要
专家领域是树形结构;欧几里得变换器并非如此,它会在深度上指数级稀释父子结构。双曲转向留下了一个未被问到的问题:不是网络要弯曲多少,而是曲率可能在何处影响梯度。位置是一种规律,而非旋钮:在可训练适配器上使用相同的几何结构会导致训练崩溃(17次训练崩溃,约220个GPU小时),但仅在损失层使用它却能成功训练,这就是HySAT(双曲结构感知训练),即仅在损失层使用双曲损失。通过构建并部署的六个专家语言模型(Llama 3.1和EXAONE 3.5;四种适配器策略;1800万个样本语料库;在约31.7万个优化器步骤中无NaN值),一个匹配的四臂消融实验分离出了保留的流形不变量,三个命题和一个引理证明了为何仅在损失层放置是稳定的,而在流形上使用适配器则不然。四个模型已投入实际应用(一个面向消费者的在线模型),两个模型开放权重,在Zenodo上有每步跟踪记录和一份包含17次事件的故障记录(知识共享许可协议CC-BY-4.0)。
英文摘要
Expert domains are trees; the Euclidean transformer is not, diluting parent-child structure exponentially at depth. The hyperbolic turn left one question unasked: not how much of a network to curve, but where curvature may touch the gradient. Placement is a law, not a knob: the same geometry on a trainable adapter collapses training (seventeen training collapses, ~220 GPU-hours), yet at the loss layer alone it trains without one -- this is HySAT (Hyperbolic Structure-Aware Training), hyperbolic losses at the loss layer only. Across six expert SLMs we constructed and deployed (Llama 3.1 and EXAONE 3.5; four adapter strategies; 18.0M-sample corpus; zero NaN over ~317K optimizer steps), a matched four-arm ablation isolates the preserved manifold invariant, and three propositions and a lemma prove why loss-only placement is stable where adapter-on-manifold is not. Four models are operationally deployed (one live, consumer-facing), two open-weight, with per-step traces and a seventeen-incident failure ledger on Zenodo (CC-BY-4.0).
Comments40 pages, 11 figures. Supplementary Information included as an ancillary file. Data and code: Zenodo, concept DOI 10.5281/zenodo.21438499 (published, CC-BY-4.0)