arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.20847cs.CLcs.SI

阅读焦虑还是阅读标签?比较微调模型与前沿模型在社交媒体上的焦虑检测

Reading Anxiety or Reading the Label? Comparing Fine-Tuned and Frontier Models for Anxiety Detection on Social Media

Cris Huynh, Arlene Pham

AI总结:

本研究比较了六种模型在Reddit焦虑检测任务上的表现,发现前沿模型F1最高(0.846),但110M参数微调编码器接近(0.831)且无API依赖;同时揭示关键词匹配混淆,微调模型词汇依赖性更高,已发表结果实为上限。

AI中文摘要:

焦虑是最常见的心理健康状况之一,人们往往在寻求临床帮助之前就已在网上写下相关内容。构建检测工具的从业者面临一个具体选择:调用前沿商业模型、在内部微调较小的模型,或部署传统分类器。我们在单一受控协议下,在留出的Reddit测试集上比较了涵盖这三种方式的六种条件。我们还发现了该任务评估方式中的一个混淆因素。在此使用的语料库中,69.3%的焦虑标注帖子包含“焦虑”一词或其变体,约为可比条件的两倍,因此分类器可以通过关键词匹配而非建模该状况的语言来获得高分。因此,我们对每个模型在原始文本和删除这些术语后的文本上各评估一次,并将差异报告为词汇依赖性。一个前沿模型在焦虑F1分数上领先(0.846),但一个110M参数的领域自适应编码器达到了0.831,且无外部API依赖,而心理健康领域预训练仅贡献了其中0.7个百分点。词汇依赖性跨度从8.6到25.4个百分点,且不随模型能力变化:LoRA微调的3B模型是测试中关键词依赖性最高的条件,甚至高于TF-IDF分类器,而前沿零样本模型依赖性最低。因此,该语料库上已发表的数字是上限,且对于此类数字通常报告的微调模型,其膨胀最为严重。

英文摘要:

Anxiety is among the most common mental health conditions, and people often write about it online well before seeking clinical help. Practitioners building detection tools face a concrete choice: call a frontier commercial model, fine-tune a smaller model in-house, or deploy a conventional classifier. We compare six conditions spanning all three on a held-out Reddit test set under a single controlled protocol. We also identify a confound in how this task is evaluated. In the corpus used here, 69.3% of anxiety-labelled posts contain the word "anxiety" or a variant, roughly twice the rate of comparable conditions, so a classifier can score well by keyword matching rather than by modelling the language of the condition. We therefore evaluate every model twice, on original text and with those terms deleted, and report the difference as lexical dependence. A frontier model leads on anxiety F1 (0.846), but a 110M-parameter domain-adapted encoder reaches 0.831 with no external API dependency, and mental-health domain pretraining accounts for only 0.7 of those points. Lexical dependence spans 8.6 to 25.4 points and does not track model capability: the LoRA fine-tuned 3B model is the most keyword-dependent condition tested, above even a TF-IDF classifier, while the frontier zero-shot model is the least. Published figures on this corpus are therefore upper bounds, and the inflation is largest for the fine-tuned models such figures typically report.

补充信息

↑