arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

高斯LoRA的认证前沿:独立先验、后验风险与预测保持平衡

Certification Frontiers for Gaussian LoRA: Independent Priors, Posterior Risk, and Prediction-Preserving Balancing

Joyanta Jyoti Mondal, Ibne Farabi Shihab

arXiv 2609.32271首次发表:更新:

AI 中文总结

本研究刻画高斯LoRA后验的PAC-Bayes认证前沿,区分先验、预测器与复杂度干预,证明仅降低复杂度无法在记录预算下认证,而Chernoff核算优于Hoeffding,矩阵平衡可显著降低KL且不改变预测。

AI 中文摘要

事后贝叶斯微调在训练好的低秩适配器周围放置高斯分布,然而校准后的后验本身并不能产生有用的泛化认证。只有当其采样预测器的损失及其与可接受先验的KL散度都较小时,这样的后验才能获得信息丰富的PAC-Bayes认证。在本研究中,我们刻画了高斯LoRA后验的认证前沿,并区分了三种干预措施:改变先验、改变随机预测器、以及仅改变其复杂度计算方式。首先,精确的各向同性KL包络消除了先验尺度,并产生一个宽度阈值,在采样前排除后验宽度,而零KL下限则识别出在测量风险界下任何复杂度降低都无法达到的目标。其次,我们在低秩因子的完整$GL(r)$对称性上以闭式最小化KL,保持每个采样的适配器乘积不变,并推导出当先验中心在独立分割上训练时所需的非中心目标。在对567种配置的小规模RoBERTa审计中,记录的64次抽取摘要意味着即使KL为零且采用Chernoff核算,认证下限也为$0.7298$,因此仅降低复杂度无法在记录预算下认证这些后验。对于受控高斯因子任务中的稳定后验,Chernoff核算在20个数据集中的20个上以1024次抽取认证风险低于$0.1$,而Hoeffding则无法认证任何数据集。在故意变形的合成秩四因子中,矩阵平衡在平均上将KL降低了$29.3\%$,超过标量平衡,且不改变任何预测。数值分割先验场景使剩余风险和复杂度预算变得明确。

英文摘要

Post-hoc Bayesian fine-tuning places Gaussians around trained low-rank adapters, yet a calibrated posterior does not by itself yield a useful generalization certificate. Such a posterior admits an informative PAC-Bayes certificate only when both the loss of its sampled predictors and its KL divergence from an admissible prior are small. In this research, we characterize this certification frontier for Gaussian LoRA posteriors and separate three interventions: changing the prior, changing the stochastic predictor, and changing only how its complexity is counted. First, an exact isotropic KL envelope eliminates the prior scale and yields a width threshold that excludes posterior widths before sampling, while a zero-KL floor identifies targets that no complexity reduction can reach at a measured risk bound. Second, we minimize KL in closed form over the full $GL(r)$ symmetry of the low-rank factors, leaving every sampled adapter product unchanged, and derive the noncentral objective required when the prior center is trained on an independent split. On a small-pool RoBERTa audit of 567 configurations, the recorded 64-draw summaries imply a certificate floor of $0.7298$ even with zero KL and Chernoff accounting, so reducing complexity alone cannot certify these posteriors at the recorded budgets. For a stable posterior in a controlled Gaussian-factor task, Chernoff accounting certifies risk below $0.1$ on 20 of 20 datasets with 1024 draws, whereas Hoeffding certifies none. On deliberately deformed synthetic rank-four factors, matrix balancing reduces KL by $29.3\%$ on average beyond scalar balancing without changing any prediction. Numerical split-prior scenarios make the remaining risk and complexity budgets explicit.

Comments32 pages, 9 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑