arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

寻找特定任务的视觉路径:跨视觉语言模型的进化块剪枝

Searching for Task-Specific Vision Paths: Evolutionary Block Pruning Across Vision-Language Models

Tarun Tomar

arXiv 2607.17052首次发表:更新:

发表机构

IIIT(国际信息技术研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究视觉语言模型中能否跳过固定预算视觉块组合,引入源平衡进化搜索并与其他方法比较,实验表明组合搜索可改进路由构建,但能力标签不能定义普遍可转移视觉路径,搜索能跨架构转移,能力专业化稳定性欠佳。

AI 中文摘要

视觉语言模型通常对每个问题都执行相同的完整视觉编码器,即使光学字符识别、计数、对象、属性和空间查询可能不需要相同的计算。我们研究了在不进行微调的情况下,是否可以跳过固定预算的视觉块组合。共享的K块路由为每个问题跳过一组恰好K个搜索到的块,而特定能力的K块策略使用已知的能力标签选择一条相同大小的路由。我们引入了一种源平衡进化搜索,并将其与独立排名、连续移除和匹配预算下的随机路由进行比较。实验使用Qwen2.5-VL-3B-Instruct、SmolVLM2-2.2B-Instruct和一个876个示例的图像不相交选择分割。搜索可跨架构转移:在SmolVLM2上,搜索到的共享四块路由比独立构建高出4.91个百分点。能力专业化不太稳定。在Qwen上,六块能力策略比共享路由高出2.17分,这是由7.10分的光学字符识别增益驱动。然而,在密封的IIIT5K上,SmolVLM2的特定光学字符识别路由比其共享路由落后13.6分。组合搜索可靠地改进了路由构建,但能力标签并不能定义普遍可转移的视觉路径。

英文摘要

Vision-language models normally execute the same complete vision encoder for every question, even when OCR, counting, object, attribute, and spatial queries may not require identical computation. We study whether fixed-budget combinations of vision blocks can be skipped without fine-tuning. A shared K-block route skips one searched set of exactly K blocks for every question, while a capability-specific K-block policy selects one same-size route using a known capability label. We introduce a source-balanced evolutionary search and compare it with independent ranking, contiguous removal, and random routes at matched budgets. Experiments use Qwen2.5-VL-3B-Instruct, SmolVLM2-2.2B-Instruct, and an 876-example image-disjoint selection split. Search transfers across architectures: on SmolVLM2, the searched shared four-block route beats independent construction by 4.91 percentage points. Capability specialization is less stable. On Qwen, the six-block capability policy beats the shared route by 2.17 points, driven by a 7.10-point OCR gain. On sealed IIIT5K, however, the SmolVLM2 OCR-specific route trails its shared route by 13.6 points. Combinatorial search reliably improves route construction, but capability labels do not define universally transferable vision pathways.

Comments14 pages, 8 figures. Code and aggregate evidence: https://github.com/TarunTomar122/vision-pathways

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑