arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

分析关于城市主义的公共话语:基于YouTube评论的主题聚类、情感分析与检索增强生成

Analyzing Public Discourse on Urbanism: Topic Clustering, Sentiment Analysis and Retrieval-Augmented Generation using YouTube Comments

Jakob Morales, Monica Hegde, Fayeq Jeelani Syed

arXiv 2609.22705首次发表:更新:

发表机构

Luddy School of Informatics, Computing and Engineering, Indiana University(印第安纳大学卢迪信息学、计算与工程学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出结合地理实体解析、主题建模、情感分析和RAG的流水线系统,分析309个北美城市的YouTube评论,并测量标准NLP组件在短非正式文本上的性能表现。

AI 中文摘要

关于城市议题——步行友好性、自行车基础设施、公共交通、住房密度和街道安全——的在线话语数量庞大但缺乏结构,现有的城市评估工具未能捕捉这些内容。我们提出了一套流水线和对话系统,结合了地理实体解析、主题建模、情感分析和检索增强生成(RAG),应用于覆盖309个北美城市的22,788个YouTube转录文本和评论片段。除了系统本身,我们的贡献在于一系列测量结果,揭示了当标准NLP组件遇到简短、非正式、地理上模糊的文本时会发生什么。一个针对Twitter调优的RoBERTa分类器在宏平均F1分数上比VADER词典基线高出12.6个百分点(0.589对0.464;McNemar检验p=0.0001),但两个模型在主导城市主义评论流量的中性类别上均表现崩溃;标注者在该类别上意见不一致(Cohen's kappa=0.53)。稠密检索在每个截断点上都优于TF-IDF基线(P@5为0.790对0.560),而视频级别的相关性代理严重低估了片段级别的精确度(人工评分下为0.660对0.94)。对于接地性评估,我们发现当将多句生成的摘要与单个简短评论进行比较时,BERTScore不可用——无论相关性如何,分数几乎持平——并表明基于ROUGE-1的接地性是一个由释义驱动的下界,而非幻觉率。这些发现超越了城市主义领域,适用于任何构建在简短用户生成文档之上的RAG系统。

英文摘要

Online discourse about urban issues - walkability, cycling infrastructure, public transit, housing density, and street safety - is voluminous but unstructured, and existing city-evaluation tools capture none of it. We present a pipeline and conversational system that combines geographic entity resolution, topic modeling, sentiment analysis, and Retrieval-Augmented Generation (RAG) over 22,788 chunks of YouTube transcripts and comments spanning 309 North American cities. Beyond the system itself, our contribution is a set of measurements about what happens when standard NLP components meet short, informal, geographically ambiguous text. A Twitter-tuned RoBERTa classifier outperforms a VADER lexicon baseline by 12.6 macro-F1 points (0.589 vs. 0.464; McNemar p = 0.0001), but both models collapse on the neutral class, which dominates urbanist comment traffic; annotators disagree on the same class (Cohen's kappa = 0.53). Dense retrieval beats a TF-IDF baseline at every cutoff (P@5 0.790 vs. 0.560), and video-level relevance proxies understate chunk-level precision by a wide margin (0.660 vs. 0.94 under human rating). For groundedness evaluation, we find BERTScore unusable when a multi-sentence generated summary is compared against a single short comment - scores are nearly flat regardless of relevance - and show that ROUGE-1-based groundedness is a paraphrase-driven lower bound rather than a hallucination rate. These findings generalize beyond the urbanist domain to any RAG system built over short user-generated documents.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑