VisionPsy-Nano:提升端侧视觉语言模型的准确性、效率与可靠性
VisionPsy-Nano: Improving Accuracy, Efficiency, and Reliability in On-Device Vision-Language Models
浏览论文内容
中文总结 AI 辅助
针对端侧视觉语言模型在准确性、效率和可靠性上的不足,提出诊断驱动的训练后方案与视觉词元策略,在0.5B模型上达到最高平均分并大幅降低延迟。
中文摘要 AI 辅助
十亿参数以下的视觉语言模型越来越适合端侧部署,但仅靠紧凑的模型规模并不能保证可用性。在手机上,此类模型生成第一个词元仍可能需要超过两分钟。端侧可用性取决于三个维度:准确性、效率和行为可靠性;标准基准忽略了第三个维度,因为其答案过短而无法暴露死循环,提示过于温和而无法探测对抗性安全性。我们提出了一种诊断驱动的训练后方案,其中教师视觉语言模型对学生模型进行压力测试,发现超出人类先验的失败模式,并将其转化为有针对性的监督和偏好对齐,用失败驱动的优化补充通用数据扩展。结合两种视觉词元策略,该方案产生了两个准确性与效率变体,并提升了行为可靠性。\textbf{\NanoFull}在17个基准上达到62.3的归一化平均值,是公开的约0.5B模型中最高的,在相同架构和词元预算下比其基础模型高出+7.4,其死循环率等于或低于最强基线。\textbf{\FlashFull}保持了61.4,同时将Pixel 9上的热启动首词元时间从138秒缩短至6.1秒(23倍)。通过同时解决这三个维度,我们推动紧凑型视觉语言模型走向实用的端侧可用性。
英文摘要
Sub-billion-parameter Vision-Language Models are increasingly viable for on-device deployment, yet compact model size alone does not guarantee usability. On a phone, such a model can still require more than two minutes to produce its first token. On-device usability depends on three axes: accuracy, efficiency, and behavioral reliability; standard benchmarks miss the third, with answers too short to expose doom loops and prompts too benign to probe adversarial safety. We introduce a diagnosis-driven post-training recipe in which a teacher VLM stress-tests the student, uncovers failure modes beyond human priors, and converts them into targeted supervision and preference alignment, supplementing generic data scaling with failure-driven optimization. Coupled with two visual-token policies, the recipe yields two accuracy-efficiency variants with improved behavioral reliability. \textbf{\NanoFull} attains a 62.3 normalized average over 17 benchmarks, the highest among openly released $\sim$0.5B models, +7.4 over its base at identical architecture and token budget, with doom-loop rates at or below the strongest baseline's. \textbf{\FlashFull} retains 61.4 while cutting warm time-to-first-token on a Pixel 9 from 138\,s to 6.1\,s (23$\times$). By jointly addressing all three axes, we move compact VLMs toward practical on-device usability.