arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.30882cs.CL

转录压缩对基于LLM的日语YouTube视频医疗 misinformation 检测的影响

Effects of Transcript Compression on LLM-based Medical Misinformation Detection in Japanese YouTube Videos

  • University of Tsukuba(筑波大学)

机构由 AI 辅助整理,请以论文原文为准。

Yuya Wake, Sho Tsugawa, Toshiyuki Amagasa

AI总结:

本研究探讨转录压缩对基于LLM的日语医疗视频虚假信息检测的影响,发现完整转录性能最佳,压缩输入增加假阴性,使虚假视频更易被误判为真实。

AI中文摘要:

大型语言模型(LLMs)越来越多地被用于评估长篇医疗视频,但其有效性可能取决于转录文本是完整提供,还是通过摘要、检索或声明筛选进行压缩。本研究探讨了此类转录压缩如何影响基于LLM的日语医疗YouTube视频真实性分类。我们比较了四种转录输入设计:完整转录、LLM生成的摘要、基于RAPTOR的检索增强生成(RAG)以及筛选(Screening),后者提取候选医疗和健康相关句子。使用74个标记为真实或虚假的长篇视频,我们评估了分类性能,并使用J-LIWC、模糊表达以及机构或技术术语分析了语言变化。完整转录基线取得了最佳性能,而所有压缩输入均增加了假阴性,意味着虚假视频更可能被误分类为真实。摘要导致性能下降最大,而筛选在压缩输入中表现最佳,但仍遗漏了许多医学相关句子。语言分析表明,这些错误并非简单地由确定性增加所解释。相反,摘要减少了情感、社交、时间、认知和对话线索,而摘要和RAG使机构和技术术语更加突出。这些发现表明,转录压缩可能使虚假视频表现为更连贯和权威的输入,从而削弱了 misinformation 检测所需的线索。

英文摘要:

Large language models (LLMs) are increasingly used to assess long-form medical videos, but their effectiveness may depend on whether transcripts are provided in full or compressed through summarization, retrieval, or claim screening. This study examines how such transcript compression affects LLM-based veracity classification of Japanese medical YouTube videos. We compare four transcript input designs: full transcripts, LLM-generated summaries, RAPTOR-based retrievalaugmented generation (RAG), and Screening, which extracts candidate medical and health-related sentences. Using 74 long-form videos labeled as Real or Fake, we evaluate classification performance and analyze linguistic changes using J-LIWC, hedge expressions, and institutional or technical terms. The full-transcript Baseline achieved the best performance, whereas all compressed inputs increased false negatives, meaning that Fake videos were more likely to be misclassified as Real. Summary caused the largest performance drop, while Screening performed best among the compressed inputs but still omitted many medically relevant sentences. Linguistic analyses showed that these errors were not explained by a simple increase in certainty. Instead, Summary reduced affective, social, temporal, cognitive, and conversational cues, while Summary and RAG made institutional and technical terms more salient. These findings suggest that transcript compression can represent Fake videos as more coherent and authoritative inputs, thereby weakening cues needed for misinformation detection

补充信息

↑