arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.12254cs.CLcs.AI

气候相关文献中社会临界点证据的自动检测与结构化:一个模块化人工智能框架

Automated Detection and Structuring of Social Tipping Point Evidence in Climate related Documents: A Modular AI Framework

  • University of Oulu(奥卢大学)
  • ITML CY(ITML CY公司)
  • PredictBy Research and Consulting SL(PredictBy研究咨询有限公司)
  • Verimpact(Verimpact公司)
  • Epsilon International Ltd(埃普西隆国际有限公司)
  • Politecnico di Milano(米兰理工大学)
  • Inspiring Futures Europe(欧洲启迪未来机构)
  • Centre for European Policy Studies(欧洲政策研究中心)

机构由 AI 辅助整理,请以论文原文为准。

Kavindu Perera, Mohammad Abaeiani, Ekaterina Gilman, Lauri Loven, Mourad Oussalah, Tassos Kanellos, Beatrice Gobbo, Dante Adami, Nicolò Ferriani, Maximiliano Ro… 展开作者

Kavindu Perera, Mohammad Abaeiani, Ekaterina Gilman, Lauri Loven, Mourad Oussalah, Tassos Kanellos, Beatrice Gobbo, Dante Adami, Nicolò Ferriani, Maximiliano Romero, Pierre Rossel, Marc Bonazountas, Christina Deligianni, Nikos Xyderis, Artur Bogucki, Lampros Argyriou, Prasasthy Balasubramanian

AI总结:

针对气候文献中社会临界点证据分散且缺乏系统发现方法的问题,提出一个模块化Transformer框架,在段落层面检测并结构化证据,实验表明其性能优于现有方法。

AI中文摘要:

气候文献的增长速度已超过评审团队的阅读速度。这一差距对于环境社会临界点这一概念尤为重要,它指的是小变化触发社会系统快速、自我强化变化的阈值。此类转变的证据通常包含在较长文档中的一两段内。因此,现有的文本挖掘工具——按主题对整篇文档进行分类或突出孤立主张——使得大量重要证据缺乏系统的发现或组织方法。本文提出一个开放且模块化的基于Transformer的框架,在段落层面检测并结构化社会临界点证据。该框架将五个组件整合为单一可部署工作流:用于分割的DistilBERT边界分割器,用于检测的迭代增强RoBERTa分类器,用于重写每个检测段落以提高清晰度的Mistral 7B模型,用于根据五个已发表的社会临界点标准对段落评分的LLaMA 3.2 3B模型,以及用于语义检索的Milvus向量存储。系统封装在Streamlit界面中,并以MinIO对象存储为后端。在由GPT-4.1标注的163段落基准和专家评审的51段落集上评估,分割器在九项指标综合得分(6.137)上超越了三种竞争方法。调优后的RoBERTa模型在完整基准上达到71.4%的准确率,Cohen's kappa为0.337;在带标签的段落上达到87.5%的准确率,kappa为0.742,优于气候专用模型和未调优的语言模型。

英文摘要:

The climate literature has grown faster than review teams can read it. That gap matters most for a concept like the environmental social tipping point, the threshold at which a small change triggers rapid, self-reinforcing change in a social system. Evidence of this kind of shift is usually contained in one or two paragraphs within a longer document. As a result, existing text mining tools-which categorize entire documents by topic or highlight isolated claims-leave an expanding set of important evidence without any systematic method for discovery or organization. This paper presents an open and modular transformer-based framework that detects and structures social tipping point evidence at the passage level. The framework joins five components into a single deployable workflow: a DistilBERT boundary splitter for segmentation, an iteratively augmented RoBERTa classifier for detection, a Mistral 7B model that rewrites each detected passage for clarity, a LLaMA 3.2 3B model that rates the passage against five published social tipping point criteria, and a Milvus vector store for semantic retrieval. The system is wrapped in a Streamlit interface backed by MinIO object storage. Evaluated on a 163-passage benchmark labelled by GPT-4.1 and a 51-passage set reviewed by experts, the splitter surpassed three competing methods on a nine-metric composite score (6.137). The tuned RoBERTa model achieved 71.4 percent accuracy with a Cohen's kappa of 0.337 on the full benchmark, and 87.5 percent accuracy with a kappa of 0.742 on passages with labels, outperforming both a climate-focused model and untuned language models.

↑