斯里兰卡议会辩论的三语主题建模
Trilingual Topic Modeling of Sri Lankan Parliamentary Debates
浏览论文内容
中文总结 AI 辅助
针对斯里兰卡议会三语辩论语料库因多语言特性难以处理的问题,提出基于LLM的文本提取结合多语言嵌入与密度聚类的端到端框架,辅以BiTopic方法,成功识别三种语言的30个宏观主题,其时间轨迹与重大国家事件吻合,且性能优于传统LDA。
中文摘要 AI 辅助
斯里兰卡议会辩论记录(Hansards)是包含僧伽罗语、泰米尔语和英语的三语语料库,其中还存在语码混合内容,但由于布局复杂的PDF文件、多语言文字系统以及黏着语形态,该语料库无法被标准自然语言处理(NLP)流程处理。本文提出一种端到端框架,通过基于大语言模型(LLM)的文本提取,再结合多语言嵌入与基于密度的聚类流程来进行主题建模。本文还探索了一种混合语义-词汇扩展方法BiTopic,以提升可解释性并恢复原本被当作噪声丢弃的演讲内容。将该框架应用于2017-2026年间的19553篇演讲,该流程成功恢复了30个宏观主题,其聚类纯度(BCP)达到0.673,这些主题的时间轨迹与2019年复活节周日袭击、2022年经济危机等重大国家事件无监督地吻合。传统的LDA因跨语言碎片化问题在该语料库上失效,而本文提出的方法能够在无监督的情况下识别三种语言的主题结构。
英文摘要
Sri Lankan parliamentary debates (Hansards) constitute a trilingual corpus of speeches in Sinhala, Tamil, and English, including code-mixed content, yet remain inaccessible to standard NLP pipelines due to layout-complex PDFs, multilingual scripts, and agglutinative morphology. We present an end-to-end framework that addresses these challenges through LLM-based text extraction followed by a multilingual embedding and density-based clustering pipeline for topic modeling. A hybrid semantic-lexical extension, BiTopic, is further explored to improve interpretability and recover speeches otherwise discarded as noise. Applied to 19,553 speeches spanning 2017-2026, the pipeline recovers 30 macro-topics achieving a cluster purity (BCP) of 0.673, whose temporal trajectories align unsupervised with major national events including the 2019 Easter Sunday attacks and the 2022 economic crisis. Traditional LDA fails on this corpus due to cross-lingual fragmentation, whereas the proposed approach successfully identifies thematic structure across all three languages without supervision.