arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.35856cs.LGcs.AI

超越关键词:利用生成式大语言模型和标签聚合对新闻文章中的经济政策不确定性进行分类

Beyond Keywords: Leveraging Generative LLMs and Label Aggregation to Classify Economic Policy Uncertainty in News Articles

Paul Trust

首次发表
浏览论文内容

中文总结 AI 辅助

本研究利用生成式大语言模型和弱监督标签聚合技术,自动分类新闻中的经济政策不确定性,克服了关键词误报和人工标注成本高的局限,并支持多标签与层次分类。

中文摘要 AI 辅助

本研究描述了在公共部门经济监测中,对大语言模型(LLMs)进行适配,以自动判断一篇文章是否讨论经济政策不确定性(EPU)并识别其具体类型。以往的研究要么依赖关键词,这通常导致较高的误报率,要么使用机器学习方法,这需要大量高质量的人工标注数据,而获取这些数据既昂贵又耗时。在本研究中,我们提出了基于弱监督技术的方法,利用生成式大语言模型通过提示生成合成标签,使该方法既经济高效又具有可扩展性。此外,我们还提出了对与EPU相关的文章进行多标签和层次分类的方法。

英文摘要

This research describes the adaptation of Large Language Models (LLMs) for economic monitoring in the public sector to automatically determine whether an article discusses Economic Policy Uncertanity (EPU) and to identify its specific type. Previous studies either rely on keywords, which often result in a high count of false positives, or use machine learning approaches that require a large number of quality human labeled data that is costly and time consuming to acquire. In this study, we propose approaches based on weak supervision techniques, using generative LLMs to create synthetic labels through prompting, making the approach both cost-effective and scalable. Additionally, we propose methods for for multi-label and hierarchical classification of articles related to EPU.

↑