arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

解开语言调查与单词计数(LIWC)的解释性和预测性作用:抑郁症相关分类中的受控替换

Disentangling the Interpretive and Predictive Roles of LIWC: Controlled Substitution in Depression-Related Classification

Hsiang-Chen Yeh, Xiutian Zhao, Aurosweta Mahapatra, Shreeram Suresh Chandra, Ryan L. Boyd, Berrak Sisman

arXiv 2607.22952首次发表:更新:

AI 中文总结

研究探讨LIWC在抑郁症相关分类中的作用,通过与三种局部替换版本比较,在多语料库中评估其对分类的影响,结果显示其在固定表示下增益有限,作为解释层有用,但结论不适用于特定架构。

AI 中文摘要

语言调查与单词计数(LIWC)提供了可审计的心理语言学类别,广泛用于解释与抑郁症相关的语言,但其在现代多模态系统中的增量预测作用仍不明确。我们在匹配的参与者级交叉验证下,对五个英文和中文抑郁症相关语料库评估LIWC。我们探讨LIWC是否能改善分类以及性能变化反映了什么。将完整的LIWC与三种局部替换版本比较:PCA旋转版本、参与者打乱版本和随机边缘版本。结果表明在固定表示上下文中,LIWC在冻结的参与者级早期融合下增益有限,且无预设对比经多重比较校正幸存。单独的SBERT校准产生更大分离,但未解决五语料库的能力限制。LIWC作为可审计的、语料库条件解释层仍有用。这些结论不适用于微调、序列感知或端到端架构。

英文摘要

Linguistic Inquiry and Word Count (LIWC) provides auditable psycholinguistic categories that are widely used to interpret depression-related language, but its incremental predictive role in modern multimodal systems remains unclear. We evaluate LIWC across five English and Chinese depression-related corpora under matched participant-level cross-validation. We ask whether LIWC improves classification and what any performance change reflects. Intact LIWC is compared with three fold-local substitutes: a PCA-rotated version that removes direct access to named category coordinates, a participant-shuffled version that preserves real LIWC profiles while breaking participant alignment, and a random-marginal version that preserves feature-wise distributions. Across multiple fixed representation contexts, the results provide limited evidence for stable LIWC gains under frozen, participant-level early fusion. None of the prespecified dataset-blocked contrasts survives multiple-comparison correction. A separate SBERT calibration produces larger observed intact-versus-shuffled and intact-versus-random separations, indicating that larger participant-aligned signals can produce correspondingly larger separations under the same procedure, while not resolving the five-corpus power limitation. LIWC remains useful as an auditable, corpus-conditioned interpretive layer. These conclusions should not be generalized to fine-tuned, sequence-aware, or end-to-end architectures.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑