V-REX:面向兽医X光片的高效专用视觉语言模型(VLM)训练
V-REX: Efficient Specialist VLM Training for Veterinary X-Rays
查看机构详情
- Vyyo AI
- Mars Petcare(玛氏宠物护理)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
该研究针对兽医X光片领域,提出高效专用VLM训练方案,无需依赖额外数据,用更少资源开发出能生成兽医X光诊断报告的模型,性能远超同类开源基础模型。
中文摘要 AI 辅助
尽管通用视觉语言模型(VLM)训练成本高昂,但人们普遍认为创建领域专家需要对规模越来越大的基础模型进行微调。我们证明,在兽医放射学领域,这一假设是错误的。通过重新思考整个VLM流程——从文本分词、预训练到接地和推理——我们表明,精心设计的工程方案能够从零开始生成优于更大规模基础模型的模型,且无需依赖任何其他数据。我们的方法引入了生成式预训练和接地的新策略,这些策略可提高训练效率,提升数据利用率和下游性能。仅使用当代通用模型一小部分的参数、数据和计算资源,我们开发出首个能够生成兽医X光片诊断报告的VLM,在该任务上的表现大幅超越开源基础模型。
英文摘要
While generalist VLMs are expensive to train, creating domain experts is widely assumed to require fine-tuning increasingly large foundation models. We show that, in veterinary radiology, this assumption is misguided. By rethinking the entire VLM pipeline - from text tokenisation and pre-training to grounding and inference - we demonstrate that careful engineering can yield models that outperform much larger foundation models from scratch, without relying on any other data. Our approach introduces new strategies for generative pre-training and grounding that improve training efficiency, increasing data utilisation and downstream performance. Using only a fraction of the parameters, data, and compute of contemporary generalist models, we develop the first VLM capable of generating diagnostic reports for veterinary radiographs, surpassing open foundation models on this task by significant margin.