发表机构
APAC Research Group, Faculty of Electrical Engineering, K. N. Toosi University of Technology(亚太研究小组,电气工程学院,伊朗德黑兰的KN Toosi科技大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对轻度认知障碍筛查中数据与诊断问题,提出用冻结DINOv2 - Small模型经特定提示令牌参数高效调优的框架,结合自适应焦点损失处理边界模糊性,在交叉验证中表现良好,优于ResViT基线。
AI 中文摘要
轻度认知障碍是认知衰退的关键早期阶段,常先于阿尔茨海默病,但从神经心理学绘图测试中自动检测它受数据稀缺、类别不平衡和临床边界附近诊断模糊性的根本限制。现有方法使用计算昂贵的完全微调混合架构,将空间可解释性后置近似而非内在模型属性。我们提出一个参数高效框架,利用冻结的DINOv2 - Small模型,通过三个特定模态可学习提示令牌适配,有119万个可训练参数,令牌在源图像补丁令牌的共享交叉注意力层中作为查询。通过注意力图直接实现空间可解释性。通过注意力模块融合任务条件嵌入以量化每个受试者的模态级重要性。为处理边界模糊性,引入适应蒙特利尔认知评估量表的焦点损失,将连续认知分数整合到训练目标、损失调制和自适应样本加权中。在分层五折交叉验证下,该架构在MCI类别F1上达到0.641,AUC为0.795,在MCI类别F1上比计算量更大的ResViT基线高出0.110。
英文摘要
Mild Cognitive Impairment is a critical early stage of cognitive decline that frequently precedes Alzheimer's disease, yet its automated detection from neuropsychological drawing tests remains fundamentally constrained by data scarcity, class imbalance, and diagnostic ambiguity near clinical boundaries. Existing methodologies attempt to bypass these constraints using computationally expensive, fully fine-tuned hybrid architectures that relegate spatial explainability to a post-hoc approximation rather than an intrinsic model property. We propose a parameter-efficient framework utilizing frozen DINOv2-Small model adapted via three modality-specific learnable prompt tokens while Operating with 1.19 million trainable parameters, each token serves as a query in a shared cross-attention layer over the source image patch tokens. Crucially, spatial explainability is achieved directly through these attention maps; as a structural consequence of the architecture. Then task-conditioned embeddings fused via an attention module to quantify modality-level importance per subject. To handle boundary ambiguity, a MoCA-adapted focal loss introduced that integrates continuous cognitive scores into the training target, loss modulation, and adaptive sample weighting, strictly generalizing standard soft-label approaches. Under stratified five-fold cross-validation, the proposed architecture yields an MCI-class F1 of 0.641 and an AUC of 0.795, outperforming the computationally heavier ResViT baseline by 0.110 in MCI-class F1.