arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.22721cs.CVcs.AI

用于精神分裂症认知康复的交互式视觉语言平台

An Interactive Vision Language Platform for Cognitive Remediation in Schizophrenia

Nassira Ait Mehdi, Milissa Temmam, Slimane Larabi

首次发表
浏览论文内容

中文总结 AI 辅助

针对精神分裂症患者认知康复任务中动作评估依赖人工观察的问题,提出基于视觉语言模型的自动化框架,通过分析视频和微调模型验证动作,收集数据集评估,有效连接物理与临床反馈,提供可扩展客观方案。

中文摘要 AI 辅助

认知康复任务常要求患者进行涉及物体操作和顺序推理的结构化动作,对精神分裂症患者至关重要。但评估这些身体动作的正确性通常依赖临床医生的人工观察,存在主观性且限制了治疗干预的可扩展性。本文提出一个基于视觉语言模型的自动化框架,用于精神分裂症康复的认知康复任务中的动作验证。该系统依赖由道路、环形交叉路口、公园和玩具车等结构化微缩场景组成的摄像头监控桌面环境。患者接收描述通过操纵玩具车执行面向目标的空间动作的音频指令。为在无需持续临床监督的情况下验证所执行动作的正确性,系统分析跟踪患者手部和玩具运动的视频馈送。一个经过微调的视觉语言模型解释记录的视频序列并生成观察活动的语义描述,从而能够根据初始文本指令对执行的动作进行高级验证。收集了一个包含4634个桌面认知康复视频场景的专用数据集来评估所提出的方法。实验结果表明,我们的专门框架有效地将低级物理遥测与高级临床反馈联系起来,为高级认知康复提供了一个可扩展且客观的解决方案。

英文摘要

Cognitive remediation tasks often require patients to perform structured actions involving object manipulation and sequential reasoning. For patients diagnosed with schizophrenia, these tasks are crucial for addressing severe cognitive deficits. However, evaluating the correctness of these physical actions generally relies on manual observation by clinicians, which introduces subjectivity and limits the scalability of therapeutic interventions. In this paper, we propose an automated framework based on Vision-Language Models for action verification in cognitive remediation tasks tailored for schizophrenia rehabilitation. The proposed system relies on a camera-monitored tabletop environment composed of structured miniature scenes including roads, a roundabout, a park, and toy vehicles. Patients receive audio instructions describing goal-oriented spatial actions to perform by manipulating a toy vehicle. These interactive physical activities are specifically designed to stimulate targeted cognitive functions, such as sustained attention, motor coordination, spatial navigation, and cognitive flexibility. To verify the correctness of the performed actions without requiring continuous clinical oversight, the system analyzes the video feed tracking the patient's hand and toy movements. A fine-tuned Vision-Language Model interprets the recorded video sequences and generates semantic descriptions of the observed activities, enabling high-level verification of the executed actions with respect to the initial textual instructions. A dedicated dataset of 4634 tabletop cognitive remediation video scenarios was collected to evaluate the proposed approach. Experimental results demonstrate that our specialized framework effectively bridges low-level physical telemetry with high-level clinical feedback, presenting a scalable and objective solution for advanced cognitive rehabilitation.

发表机构

  • Computer Science Faculty, USTHB University(USTHB大学计算机科学学院)
  • RIIMA Laboratory(RIIMA实验室)

机构由 AI 辅助整理,请以论文原文为准。

↑