arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MSA-CITE:面向固定预算小模型推理的协同适配LoRA专家生态

MSA-CITE: A Co-Adapted LoRA Specialist Ecology for Fixed-Budget Small-Model Inference

Ruitong Li, Binjie Guo, Aisheng Mo, Guowei Su, Jie Li, Ru Zhang

arXiv 2609.26217首次发表:更新:

发表机构

The University of Hong Kong; Zhejiang University(香港大学; 浙江大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出MSA-CITE方法,利用多个训练后LoRA分支作为可组合资产,在固定生成预算下通过等价类分组与校准先验投票,提升小模型推理准确率,实验显示优于最强单分支基线。

AI 中文摘要

紧凑型语言模型通常通过保留单个训练后检查点并重复采样来部署。在这项工作中,我们挑战了这一做法,将多个被丢弃的检查点视为可组合的部署资产。从单个Qwen3-4B骨干网络出发,我们保留了四个冻结的LoRA分支,每个分支来自不同的训练后轨迹。我们没有从一个分支生成四个样本,而是通过从每个分支采样一个完成结果来分配固定的四样本生成预算。我们的方法,即多路径专家适配与校准推理时证据(MSA-CITE),通过将终端答案分组为等价类、使用校准得出的源先验之和为每个类打分,并在确定性平局打破规则下选择一个代表来处理生成的组合。读出阶段不从评估结果中学习,也不引入额外的生成、验证器或重排序步骤。在200个保留的数学题目上,四路径组合达到了65.5%的准确率,而最强的单分支基线为62.0%。在100个主题不相交的偏移上,它达到了42.0%,而基线为40.0%。在分布内条件下,相对于同质SFT和在线OPD重复的改进是稳健的;而相对于最强基线和偏移条件下的结果并不具有决定性。我们的发现提供了一个狭窄但具体的贡献:训练后的分支,即使没有协同训练,也能在部署中产生集体效益。

英文摘要

Compact language models are typically deployed by retaining a single post-training checkpoint and sampling it repeatedly. In this work, we challenge this practice by treating multiple discarded checkpoints as composable assets for deployment. Starting from a single Qwen3-4B backbone, we preserve four frozen LoRA branches, each derived from a different post-training trajectory. stead of drawing four generations from one branch, we allocate a fixed four-generation budget by sampling one completion from each branch. Our method, Multi-path Specialist Adaptation with Calibrated Inference-Time Evidence (MSA-CITE), processes the resulting portfolio by grouping terminal answers into equivalence classes, scoring each class via summed calibration-derived source priors, and selecting a representative under deterministic tie-breaking rules. The readout stage does not learn from evaluation results, nor does it introduce additional generations, verifiers, or reranking steps. On 200 held-out mathematics items, the four-path portfolio achieves 65.5% accuracy, compared with 62.0% for the strongest single-branch baseline. On a 100-item subject-disjoint shift, it attains 42.0% versus 40.0%. Under in-distribution conditions, the improvements over homogeneous SFT and Online-OPD repetition are robust; results against the strongest baseline and under shifted conditions are not conclusive. Our findings offer a narrow but concrete contribution: post-training branches, even without co-training, can be collectively beneficial for deployment.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑