AI 中文总结
研究旨在解决端侧视觉语言模型在准确性和效率间的矛盾,提出StepX-Edge模型,通过架构、训练和部署协同设计,在多项任务中表现出色,参数少且运行稳定,还将开源相关数据、方法和管道。
AI 中文摘要
在终端设备上部署具有完整用户界面理解能力的视觉语言模型长期以来一直困于准确性和效率之间:一方面是光学字符识别、屏幕理解、视觉问答和元素定位的准确性要求;另一方面是移动芯片严格的计算、内存和功率预算。现有工作要么顾此失彼,要么止于模拟而无真实设备验证。我们提出了StepX-Edge,一个具有0.9B参数的端侧用户界面视觉语言模型,通过架构、训练和部署的三层协同设计解决了这种矛盾。在架构上,UI感知分层视觉编码(ULVE)和渐进维度投影(PDP)连接器针对屏幕的极端宽高比和细粒度感知,同时全程标准的全注意力确保与主流移动NPU算子的原生兼容性。对于训练,五步的StepX-Curriculum框架围绕我们对用户界面子任务间相互促进作用的观察而设计,使所有四项能力在严格的参数预算下协同增长而非相互干扰。对于部署,基于模块的差异化两阶段从参数张量量化到量化感知训练的量化方案使量化后精度损失控制在1%以内。StepX-Edge在<=1B模型中实现了最强的整体用户界面理解能力,在屏幕问答(88.76 F1)和中文OCRBench v2(57.25)上超越了所有2B-2.3B基线,并在RefCOCO(92.0%)和OCRBench v1(831)上以少得多的参数与1.3B-2.3B通用视觉语言模型匹配。经过W4A16+KV8量化后,该模型在骁龙8 Gen5设备上稳定运行,推理时间约0.84秒,解码速度98词元/秒,峰值内存1.4GB。我们将开源训练数据、完整训练方法和量化部署管道。
英文摘要
Deploying a vision-language model with full UI understanding on end devices has long been trapped between accuracy and efficiency: on one side is the accuracy bar for OCR, screen understanding, visual question answering, and element grounding; on the other is the strict compute, memory, and power budget of mobile chips. Existing work either trades one for the other, or stops at simulation without real-device validation. We present StepX-Edge, a 0.9B-parameter on-device UI vision-language model that resolves this tension through three-layer co-design of architecture, training, and deployment. Architecturally, UI-aware Layered Visual Encoding (ULVE) and a Progressive Dimensionality Projection (PDP) connector target the extreme aspect ratios and fine-grained perception of screens, while standard full attention throughout ensures native compatibility with mainstream mobile NPU operators. For training, the five-stage StepX-Curriculum framework is designed around our observation of mutual-promotion effects among UI subtasks, so that all four capabilities grow synergistically under a tight parameter budget rather than interfering. For deployment, a module-wise differentiated two-stage PTQ-to-QAT quantization scheme keeps the post-quantization accuracy loss within 1%. StepX-Edge achieves the strongest overall UI understanding among <=1B models, surpassing all 2B-2.3B baselines on ScreenQA (88.76 F1) and Chinese OCRBench v2 (57.25), and matching 1.3B-2.3B general VLMs on RefCOCO (92.0%) and OCRBench v1 (831) with far fewer parameters. After W4A16+KV8 quantization, the model runs stably on Snapdragon 8 Gen5 devices with ~0.84 s TTFT, 98 tok/s decode, and 1.4 GB peak memory. We will open-source the training data, the full training recipe, and the quantization deployment pipeline.