arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

利用编译:小型视觉语言模型应保留哪些决策?

Harness Compilation: Which Decisions Should a Small Vision-Language Model Keep?

Minhao Fan, Yinyi Liu, Jiayu Zhao, Zihan Teng, Song Chen, Weichen Liu

arXiv 2610.11231首次发表:更新:

发表机构

College of Computing and Data Science, Nanyang Technological University; School of Microelectronics, University of Science and Technology of China(南洋理工大学计算与数据科学学院; 中国科学技术大学微电子学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出Harness Compilation(HC),通过调整小型视觉语言模型与外部harness的工作分配,在7项视觉问答任务中使小型VLM得分提升9.9-23.9点,且优于仅输出答案的LoRA方法。

AI 中文摘要

小型视觉语言模型(VLM)或许能够读取外部证据,但难以获取这些证据。我们提出Harness Compilation(HC),这是一种离线程序,用于调整冻结的小型VLM与其外部harness之间的工作分配。大型教师模型利用学生的执行轨迹来修改可复用的内容与控制逻辑,同时用单独的验证集选择部署的harness。部署过程既无需更新权重,也无需调用教师模型。在7项视觉问答场景中,学生模型参数最多为90亿,HC使原始学生模型的得分提升了9.9至23.9个百分点,每个场景平均进行3次独立构建。对5种运行时决策类型(调用、选择、参数生成、证据整合及弃权(不执行))的干预实验揭示了这种分配的重要性:请求证据和生成开放式查询成本较高,而受限选择与读取提供的文本可保留学生模型的有效工作。事实卡片对所有10个被评估的学生模型均有益,但决策策略的迁移效果参差不齐。当迁移的接口不再适配新学生模型时,对其重新编译会有帮助。仅用100个训练样本,HC在3项任务上就优于仅输出答案的LoRA方法。更大的训练预算可达到或超越固定harness的效果,而将两者结合能使SlideVQA的表现优于单独使用其中任何一种。这些发现支持从可测量的学生行为出发分配工作,而非统一移除决策。

英文摘要

Small vision-language models may be able to read external evidence yet struggle to obtain it. We introduce Harness Compilation (HC), an offline procedure that adapts the division of work between a frozen small VLM and its external harness. A large teacher uses student execution traces to revise reusable content and control, while a separate validation set selects the deployed harness. Deployment requires neither weight updates nor teacher calls. Across seven visual question-answering settings with students of at most 9B parameters, HC improves scores over bare students by 9.9-23.9 points, averaged over three independent builds per setting. Interventions on five runtime decision types (invocation, selection, argument generation, evidence integration and abstention) show why this allocation matters: requesting evidence and generating open queries can be costly, whereas bounded choices and reading supplied text can remain useful student work. Fact cards benefit all ten evaluated students, but decision policies transfer unevenly. Recompilation for a new student model helps when the transferred interface no longer fits the student. With 100 practice items, HC exceeds answer-only LoRA on three tasks. Larger training budgets can match or surpass a fixed harness, while combining the two improves SlideVQA beyond either alone. These findings support allocating work from measured student behavior rather than uniformly removing decisions.

Comments58 pages, 10 figures, including appendices

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑