发表机构
Southern Illinois University; Utah Valley University; Jazan University(南伊利诺伊大学; 犹他谷大学; 吉赞大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
VetClaw是用于兽医疾病筛查的边缘云多模态智能系统,用相机模块采集图像等并发送至视觉语言模型分类。它分离智能体交互与工作流编排,超越静态图像分类,症状引导和多模态输入提升性能,将静态模型转变为多功能安全感知系统。
AI 中文摘要
我们展示了VetClaw,一种用于早期兽医疾病筛查的边缘云多模态智能系统。VetClaw使用相机模块作为边缘传感设备,并将捕获的图像以及可选的症状描述发送到服务器托管的视觉语言模型进行零样本疾病分类。该系统将智能体交互与工作流编排分离:OpenClaw在边缘设备上提供调度、工具访问、用户交互和通知服务,而LangGraph管理有状态的筛查工作流,包括输入验证、图像传输、模型调用、安全检查、条件路由、故障处理和结构化日志记录。这种设计超越了静态图像分类,使系统能够收集视觉证据、调用外部模型、应用确定性安全规则并生成诊断支持警报。结果表明,仅图像的VLM预测仍然有限,而症状引导和多模态输入可提高零样本分类性能。因此,VetClaw将静态预测模型转变为一个能够使用工具、管理工作流、处理故障并升级不确定病例的协调的、安全感知系统。
英文摘要
We present VetClaw, an edge-cloud multimodal agentic system for early veterinary disease screening. VetClaw uses a camera module as an edge sensing device and sends captured images, together with optional symptom descriptions, to a server-hosted vision-language model for zero-shot disease classification. The system separates agent interaction from workflow orchestration: OpenClaw provides scheduling, tool access, user interaction, and notification services on the edge device, while LangGraph manages the stateful screening workflow, including input validation, image transmission, model invocation, safety checks, conditional routing, failure handling, and structured logging. This design moves beyond static image classification by enabling the system to collect visual evidence, invoke external models, apply deterministic safety rules, and generate diagnostic-support alerts. Results show that image-only VLM prediction remains limited, whereas symptom-guided and multimodal inputs improve zero-shot classification performance. Thus, VetClaw transforms a static prediction model into a coordinated, safety-aware system that can use tools, manage workflows, handle failures, and escalate uncertain cases.