arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

测试时的AI辅助AI:通过 harness 实现强模型到弱模型的能力迁移

AI4AI at Test-Time: Strong-to-Weak Capability Transfer via Harnesses

Cheng Qian, Wenting Zhao, Liangwei Yang, Heng Wang, Jielin Qiu, Heng Ji, Silvio Savarese, Huan Wang, Shelby Heinecke

arXiv 2608.12307首次发表:更新:

发表机构

Salesforce AI Research; University of Illinois Urbana-Champaign(Salesforce人工智能研究院; 伊利诺伊大学厄巴纳-香槟分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出在测试时通过构建 harness,无需重新训练即可将强模型的认知结构迁移给弱模型,使目标模型在四个心理理论基准上的平均性能从0.49提升至0.91,为模型能力迁移提供了新的思路。

AI 中文摘要

近期关于知识蒸馏的研究通常通过教师强制、在线策略蒸馏等训练时方法,更新小型模型的参数以将大模型的能力迁移给小型模型。本文研究这种迁移是否可在测试时实现,探讨强模型到弱模型的支架式辅助:更强的构建者模型能否构建测试时的 harness,帮助较弱的目标模型无需任何参数更新即可更可靠地解决任务。研究采用四个代表性的心理理论基准,每个构建者模型使用5%的数据作为验证集,在多轮迭代中优化其 harness,之后将最终的 harness 在完整测试集上评估。实验表明,这种测试时能力迁移形式非常有效,使目标模型的平均性能几乎翻倍,从0.49提升至0.91。分析显示,性能提升主要源于将不稳定的模型推理卸载到确定性代码、基准特定路由及严格答案格式执行,而非鼓励目标模型进行更广泛的推理或更多采样。进一步研究发现,构建者模型的推理工作量会单调提升 harness 质量;与构建者模型自身能力相比,平台效应较小;较弱的目标模型获得的提升最大。这些结果表明,测试时 harness 设计是传统训练时知识蒸馏的重要补充,使强模型无需重新训练即可将认知结构迁移给弱模型。

英文摘要

Recent work on distillation transfers the capabilities of large models to smaller ones often by updating the latter's parameters, through teacher forcing, on-policy distillation, and related training-time methods. In this paper, we ask whether such transfer can instead occur at test time. We study strong-to-weak scaffolding: whether a stronger builder model can construct inference-time harnesses that help a weaker target model solve tasks more reliably without any parameter updates. Using four representative Theory-of-Mind benchmarks, each builder model uses 5% of the data as a validation set to iteratively refine its harness over multiple rounds, after which the finalized harness is evaluated on the full test set. Empirically, this form of test-time capability transfer is highly effective, nearly doubling average target-model performance from 0.49 to 0.91. Our analysis shows that the gains come primarily from offloading unstable model reasoning into deterministic code, benchmark-specific routing, and strict answer-format enforcement, rather than from encouraging the target model to reason more extensively or sample more broadly. We further find that builder-model reasoning effort improves harness quality monotonically, platform effects are modest relative to the builder model's own capability, and weaker target models receive the largest gains. These results suggest that inference-time harness design is an important complement to conventional training-time distillation, enabling strong models to transfer cognitive structure to weaker models without retraining.

Comments23 Pages, 12 Figures, 6 Tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑