arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.22070cs.CE

SoulGard-VL-2B:用于边缘端猫行为理解的视觉语言模型

Catellect-VL-2B: A Vision-Language Model for Edge-Based Feline Behavior Understanding

  • Catellect Research(Catellect研究院)

机构由 AI 辅助整理,请以论文原文为准。

YuHang Wu, HaoXian Liu, JunYi Wang, Jia Tao

AI总结:

本研究提出基于Qwen3-VL-2B后训练得到的边缘端VLM SoulGard-VL-2B,结合多阶段后训练方案,在RK3576芯片上实现高准确率与低延迟,可生成猫行为JSON结构化输出,提升了动物VLM的实用性。

AI中文摘要:

猫行为理解任务要求模型能够识别细微视觉线索、保证行为解释可审计,且支持低延迟、隐私敏感的部署场景。直接对通用视觉语言模型(VLM)进行提示不适用于该场景:这类模型不会先报告耳朵位置、尾巴姿态等可见证据,而是可能直接输出“放松”“害怕”“疼痛”等标签,导致输出难以验证,且与边缘端使用需求不匹配——边缘端更适合紧凑的JSON输出,而非冗长的自由格式解释。我们提出SoulGard-VL-2B,一款用于猫行为理解的边缘端VLM,可生成猫行为的JSON格式结构化输出。SoulGard-VL-2B基于Qwen3-VL-2B在SoulGardBench上进行后训练得到,SoulGardBench是我们构建的包含4万样本的图像-行为标注数据集,其中约3.8万为阶段特定训练实例,2千为保留测试集。多阶段后训练方案结合了自然语言行为预热、场感知加权(FAW)监督微调以及紧凑行为序列化。实验结果显示,配备紧凑输出序列化的SoulGard-VL-2B在行为场宏准确率达到80.62%,在RK3576边缘芯片上部署时,相较于相同参数规模的全JSON基线实现了2.51倍的加速,适用于边缘部署。我们还构建了包含3千条条目的猫行为知识库,将结构化行为字段映射到情绪和意图概念,以实现基于证据的解释。综上,这些结果表明SoulGard-VL-2B可提升以动物为中心的VLM的准确性、可审计性和可部署性。

英文摘要:

The task of Feline Behavior Understanding requires models that can identify subtle visual cues, keep behavior interpretations auditable, and support low-latency, privacy-sensitive deployment. Directly prompting general Vision-Language Models (VLMs) is poorly suited to this setting: instead of first reporting visible evidence such as ear position and tail posture, they may jump directly to labels such as relaxed, afraid, or in pain. This makes the output difficult to verify and poorly aligned with edge-based use, where compact JSON outputs are preferable to long free-form explanations. We present Catellect-VL-2B, an edge-based VLM for Feline Behavior Understanding that generates JSON-formatted Structured Output for feline behavior. Catellect-VL-2B is post-trained from Qwen3-VL-2B on SoulGardBench, our 40K-sample image-behavior annotation dataset with approximately 38K stage-specific training instances and a 2K held-out test set. The multi-phase Post-Training recipe combines natural-language behavior warmup, Field-Aware Weighted (FAW) supervised fine-tuning, and compact behavior serialization. Experimental results show that SoulGard-VL-2B equipped with compact output serialization achieves 80.62 percent behavior-field macro accuracy and delivers a 2.51-fold speedup over its full-JSON baseline of identical parameter size when deployed on the RK3576 edge chip, making it suitable for edge deployment. We further build a 3K-entry feline behavior knowledge base that maps structured behavior fields to emotion and intent concepts for evidence-grounded interpretation. Together, these results show that SoulGard-VL-2B can make animal-centered VLMs more accurate, auditable, and deployable.

↑