arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.11591cs.CV

用于神经架构合成的基于相似度引导的大语言模型课程微调

Similarity-Guided Curriculum Fine-Tuning of LLMs for Neural Architecture Synthesis

  • Computer Vision Lab, CAIDAS & IFI, University of Würzburg(计算机视觉实验室、CAIDAS与IFI,维尔茨堡大学)

机构由 AI 辅助整理,请以论文原文为准。

Anujaya Vijayakumar, Radu Timofte, Dmitry Ignatov

AI总结:

研究为基于大语言模型的神经架构搜索构建基于MinHash的相似度调度框架,通过对源代码分片划分相似度带进行渐进式课程微调,在CIFAR-10等实验中评估,揭示了接口修复和课程调度针对不同失败模式,还观察了跨数据集转移情况。

AI中文摘要:

介绍了一种基于MinHash的相似度调度框架,该框架为基于大语言模型的神经架构搜索(NAS)构建渐进式课程。使用基于128排列的MinHash签名对归一化的7-gram源代码分片,将参考池划分为相似度带,并按架构异质性递增呈现,每个阶段的最佳LoRA适配器累积合并到主干中。在CIFAR-10图像分类的LEMUR基准测试中,在OlympicCoder-7B上评估该框架,每个epoch生成15个候选架构,共六个渐进式微调步骤。该课程在高相似度水平下无需后处理修复即可达到60%的峰值成功率。在最具多样性的课程与基础模型、有与无部分接口修复的2*2消融实验表明,无修复时基础模型(峰值成功率47%)大幅优于课程模型(7%),而添加部分修复两者均达到53%。这种模式与合并级权重漂移逐渐消除评估器接口先验一致,表明接口修复和课程调度针对不同的失败模式。还报告了在SVHN上的跨数据集转移观察结果,直接生成基础模型而无课程预热时,峰值成功率为27%,精度(60.5%)远低于CIFAR-10等效情况,这与unq-family锚定架构合成难度增加一致。

英文摘要:

Introduce a MinHash-based similarity scheduling framework that constructs a progressive curriculum over neural architecture code for LLM-based neural architecture search (NAS). Using 128-permutation MinHash signatures over normalised 7-gram source code shingles, we partition the reference pool into similarity bands and present them in increasing architectural heterogeneity, with the best LoRA adapter from each stage merged cumulatively into the backbone. We evaluate the framework on OlympicCoder-7B within the LEMUR benchmark on CIFAR-10 image classification, generating N =15 candidate architectures per epoch across six progressive fine-tuning steps. The curriculum achieves 60% peak success rate at the high-similarity level without post-processing repair. A 2*2 ablation at the most diverse level curriculum versus base model, with versus without partial interface repair reveals that without repair the base model (47% peak SR) substantially outperforms the curriculum model (7% SR), while adding partial repair brings both to 53% SR. This pattern is consistent with merge-level weight drift progressively erasing evaluator-interface priors, and suggests that interface repair and curriculum scheduling target distinct failure modes. We further report a cross-dataset transfer observation on SVHN, where direct base-model generation without curriculum warmup yields 27% peak SR at substantially lower accuracy (60.5%) than the CIFAR-10 equivalent, consistent with the increased synthesis difficulty of the unq-family anchor architecture.

↑