VIS-DICT:用于社交网络抑郁症检测中缺失模态填补的视觉词典
VIS-DICT: A Visual Dictionary for Missing Modality Imputation in Social Network Depression Detection
- University of Tehran(德黑兰大学)
- University of Sistan and Baluchestan(锡斯坦和俾路支斯坦大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出Vis-Dict,一种基于词典的零参数方法,通过词语与平均图像向量的关联填补缺失视觉特征,在抑郁症检测中达到与生成网络相当的性能(F1=0.9454,ROC-AUC=0.9890)。
AI中文摘要:
追踪社交媒体帖子有助于发现抑郁症的早期迹象。近期研究表明,结合文本和图像进行抑郁症检测比仅使用文本效果更好。然而,许多社交媒体帖子没有图像,这使得多模态模型难以应用。现有的大多数方法使用检索或生成模型来填补缺失图像,但这些模型需要额外训练。在本文中,我们提出了Vis-Dict,一种基于词典的方法,通过将词语与完整训练帖子中的平均图像向量相关联来构建缺失的视觉特征。这些估计的视觉特征随后与文本结合,以跟踪用户行为随时间的变化。我们在一个社交媒体数据集上测试了Vis-Dict,使用最多512条帖子的用户时间线,并将其与其他缺失数据处理方法进行了比较。结果表明,Vis-Dict的性能与生成网络相当,达到了0.9454的F1分数和0.9890的ROC-AUC。最重要的是,Vis-Dict在图像生成方面实现了零可训练参数的情况下取得了这一强劲性能。这些发现表明,将词语直接与视觉特征相关联是处理抑郁症检测系统中缺失图像的一种有效且实用的方法。
英文摘要:
Tracking social media posts can help spot early signs of depression. Recent studies show that combining text and images works better for detecting depression than using text alone. However, many social media posts do not have images, which makes it hard to use multimodal models. Most existing methods fill in missing images using retrieval or generative models that need extra training. In this paper, we introduce Vis-Dict, a dictionary-based method that builds missing visual features by linking words to average image vectors from complete training posts. These estimated visual features are then combined with text to track changes in user behavior over time. We tested Vis-Dict on a social media dataset using user timelines of up to 512 posts and compared it with other missing-data methods. The results show that Vis-Dict performs on par with generative networks, reaching an F1-score of 0.9454 and an ROC-AUC of 0.9890. Most importantly, Vis-Dict achieves this strong performance with zero trainable parameters for image generation. These findings show that directly connecting words to visual features is an effective and practical way to handle missing images in depression detection systems.