发表机构
College of Computing and Informatics; The University of North Carolina at Charlotte; Department of Public Health and Health Administration(计算与信息学院; 北卡罗来纳大学夏洛特分校; 公共卫生与健康管理系)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对社交媒体短文本主题模型缺乏分配感知评估的问题,提出文档-主题对齐度量(DoTA),在公共卫生数据集上验证其与传统度量的互补性及与人类评估的一致性。
AI 中文摘要
主题模型被广泛用于分析公共卫生相关的社交媒体短文本,然而其评估仍主要依赖于仅关注生成主题本身的度量标准。目前缺乏能够定量评估所分配主题是否有效代表相应短文本帖子的度量方法。我们提出了文档-主题对齐度量(DoTA),这是一个感知分配的评估框架,包含衡量文档(帖子)与其分配主题之间语义对齐的度量标准。我们还引入了基于边际和判别性的变体,以捕捉主题分配的置信度和可区分性。我们在来自X平台的三个公共卫生相关社交媒体数据集上,对五种主题模型评估了DoTA,并将DoTA度量与传统基于主题的度量进行了比较。结果表明,DoTA提供了互补的评估线索,并与人类评估有意义地保持一致。这些发现确立了感知分配评估的必要性,并证明加入DoTA能够对短文本主题建模性能进行更全面且具有实际意义的评估。
英文摘要
Topic models are widely used to analyze public health-related social media short texts, yet their evaluation remains dominated by metrics that focus entirely on generated topics alone. There is a lack of metrics that quantitatively assess whether assigned topics meaningfully represent the corresponding short-text posts. We propose Document-Topic Alignment metrics (DoTA), an assignment-aware evaluation framework comprising metrics that measure semantic alignment between documents (posts) and their assigned topics. We also introduce margin-based and discriminative variants that capture topic assignment confidence and distinguishability. We evaluate DoTA across five topic models on three public health-related social media datasets from X and compare DoTA metrics with conventional topic-based metrics. Results show that DoTA provides complementary evaluation cues and aligns meaningfully with human evaluations. These findings establish the need for assignment-aware evaluation and demonstrate that the addition of DoTA enables a more comprehensive and practically meaningful evaluation for assessing short-text topic modeling performance.
CommentsAccepted for publication in the Proceedings of the 60th Hawaii International Conference on System Sciences (HICSS 2027)