arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.31298cs.CVcs.AI

UniAR:一种由多视角提示学习增强的自闭症识别统一框架

UniAR: A Unified Framework for Autism Recognition Enhanced by Multi-View Prompt Learning

Lei Xin, Zeheng Wang, Jiayin Zhu, Shihong Huang, Fanhu Zeng, Changjiang Jiang, Dengbo He, Yutao Yue, Zhenglun Kong

首次发表
浏览论文内容

中文总结 AI 辅助

针对自闭症识别中文本数据稀缺问题,提出UniAR统一框架,利用多粒度提示学习生成层次化诊断描述,并通过混合专家多尺度对齐模块融合视觉与语义,在四个基准上显著提升准确率。

中文摘要 AI 辅助

自闭症谱系障碍(ASD)是一种复杂的神经发育障碍,早期准确诊断对于改善长期发育结果至关重要。然而,现有的ASD识别方法常受限于诊断文本数据的稀缺,迫使它们主要依赖视觉分析,从而限制了其对临床上有意义的语义推理进行建模的能力。为应对这一挑战,我们提出了UniAR,一个由多粒度提示学习增强的统一框架,用于在异构数据变化下进行稳健的ASD识别。具体而言,UniAR利用大型多模态模型生成词、短语和句子级别的层次化诊断描述,弥补配对临床报告的缺失。为了将生成的语义与视觉证据对齐,我们进一步设计了一种基于混合专家(Mixture-of-Experts)的多尺度对齐模块,该模块动态地将向量量化的视觉原型与相应粒度下的语义表示进行匹配。在涵盖脑MRI和面部表情场景的四个基准上的广泛实验表明,UniAR持续优于现有最先进方法,在MRI基准上平均准确率达到75.9%,在面部基准上平均准确率达到91.6%,同时相比基线,在MRI基准上的平均准确率提高了1.5个百分点,在面部基准上的平均准确率提高了1.2个百分点。这些结果表明,UniAR为语义稀缺条件下的ASD筛查提供了一个稳健且可解释的框架。

英文摘要

Autism Spectrum Disorder (ASD) is a complex neurodevelopmental disorder for which early and accurate diagnosis is critical to improving long-term developmental outcomes. However, existing ASD recognition methods are often constrained by the scarcity of diagnostic text data, forcing them to rely mainly on visual analysis and limiting their ability to model clinically meaningful semantic reasoning. To address this challenge, we propose UniAR, a unified framework enhanced by multi-granularity prompt learning for robust ASD recognition under heterogeneous data variations. Specifically, UniAR leverages a large multimodal model to generate hierarchical diagnostic descriptions at the word, phrase, and sentence levels, compensating for the lack of paired clinical reports. To align the generated semantics with visual evidence, we further design a Mixture-of-Experts-based Multi-Scale Alignment Module, which dynamically matches vector-quantized visual prototypes with semantic representations at corresponding granularities. Extensive experiments on four benchmarks covering brain MRI and facial expression scenarios show that UniAR consistently outperforms existing state-of-the-art methods, achieving average accuracies of 75.9\% on MRI benchmarks and 91.6\% on facial benchmarks, while improving average Accuracy on MRI benchmarks by 1.5 percentage points and average Accuracy on facial benchmarks by 1.2 percentage points over baselines. These results demonstrate that UniAR offers a robust and interpretable framework for ASD screening under semantic scarcity.

发表机构

  • Wuhan University(武汉大学)
  • Northeast Normal University(东北师范大学)
  • The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
  • Nanjing Agricultural University(南京农业大学)
  • Imperial College London(帝国理工学院)
  • The Hong Kong University of Science and Technology(香港科技大学)
  • Harvard University(哈佛大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑