arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.12677cs.AIcs.CV

自然语言理解在基于多模态视频的登革热诊断中的作用

The Role of Natural Language Understanding in Multimodal Video-Based Dengue Diagnosis

Danial Sharifrazi, Saadat Behzadi, Julakha Jahan Jui, Mojtaba Mohammadi, Nouman Javed, Roohallah Alizadehsani, Prasad N. Paradkar, Asim Bhatti

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出基于YOLO和CLIP的视觉-语言框架,从视频中分类未受感染与DENV2感染的蚊子,帧级准确率达98.54%、灵敏度达99.91%,证实该框架可用于分析感染相关生物行为。

中文摘要 AI 辅助

从视频数据中检测蚊子受感染相关的行为变化颇具挑战,因为蚊子体型小、移动快速且无规律,还会受背景、光照和阴影等环境因素影响,难以进行可靠的特征提取。本研究提出了一种基于YOLO和对比语言-图像预训练(CLIP)的视觉-语言框架,用于对未受感染及登革病毒2型(DENV2)感染的蚊子飞行帧进行分类。首先使用YOLO将蚊子区域从背景中分离,随后将从视频帧提取的视觉特征与具有生物学意义的文本提示在共享嵌入空间中对齐。该多模态模型通过监督双向对比学习进行微调,并通过基于帧级图像-文本相似度的分类进行评估。结果显示,所提方法在帧级达到了98.54%的准确率和99.91%的灵敏度;对帧级信息进行时间聚合后,模型在视频级实现了完整性能。消融实验结果表明,微调及基于CLIP的表示对该领域至关重要,而文本分支提供了语义层面的图像-文本对齐,并非优于仅视觉模型的准确率优势。这些发现表明,视觉-语言模型可为分析视频数据中与感染相关的生物行为提供有用框架。

英文摘要

Detecting infection-related behavioral changes in mosquitoes from video data is challenging because mosquitoes are small, move rapidly and irregularly, and are affected by environmental factors such as background, lighting, and shadows, which can make reliable feature extraction difficult. In this study, a YOLO- and Contrastive Language-Image Pre-training (CLIP)-based vision-language framework is proposed to classify mosquito flight frames of uninfected and Dengue virus serotype 2 (DENV2)-infected mosquitoes. First, YOLO is used to isolate mosquito regions from the background. Then, visual features extracted from video frames are aligned with biologically meaningful textual prompts in a shared embedding space. The multimodal model was fine-tuned using supervised bidirectional contrastive learning and evaluated through frame-level image-text similarity-based classification. The results show that the proposed method achieved 98.54% accuracy and 99.91% sensitivity at the frame level. After temporal aggregation of frame-level information, the model achieved complete video-level performance. The ablation results showed that fine-tuning and CLIP-based representations were essential for this domain, while the textual branch provided semantic image-text alignment rather than an accuracy advantage over the vision-only model. These findings suggest that vision-language models can provide a useful framework for analyzing infection-related biological behaviors from video data.

↑