发表机构
University of Toronto(多伦多大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出一种利用网络数据和大语言模型量化组织环境行动的可扩展框架,通过对比三种检测方法,发现直接LLM分类覆盖率最高,可适用于其他机构。
AI 中文摘要
从公开可获取的网络内容中量化组织的环境行动仍然是一个具有挑战性的环境数据科学问题,因为相关信息可能分散在多个网页中,并且主要通过非结构化文本进行传达。我们提出了一种可扩展的计算框架,用于将组织网络内容转化为结构化的环境行动度量,并以美国犹太教会众为例进行了演示。我们通过整合多个地理空间、知识库、目录及人工审核来源,构建了一个包含4,964个会众的全国数据库。其中,2,657个会众拥有活跃网站并成功爬取,生成了包含154,454个网页的语料库。我们比较了三种检测环境行动的方法:关键词检索后接大语言模型(LLM)分类、语义向量检索后接LLM分类,以及无预先检索的直接LLM分类。与专家人工审核者的一致性方面,关键词检索的一致性最低(κ=0.26),语义向量检索的一致性较高(κ=0.42),直接LLM分类的一致性相近(κ=0.40)。尽管语义检索实现了最高的一致性,但其检索召回率为0.87,表明在分类前丢失了部分相关内容。应用于完整语料库时,直接LLM分类在1,398个会众(53%)中识别出至少一项环境行动,覆盖率高于两种基于检索的方法。这些结果表明,预先检索可以降低计算成本,但可能在分类器处理之前排除相关信息。该框架提供了一种可复现的方法,用于从非结构化网络内容中提取组织层面的环境信息,并可适用于其他机构。
英文摘要
Quantifying organizational environmental action from publicly available web content remains a challenging environmental data science problem because relevant information can be dispersed across multiple webpages and is primarily communicated through unstructured text. We present a scalable computational framework for transforming organizational web content into structured measures of environmental action and demonstrate the approach using Jewish congregations in the United States. We constructed a national database of 4,964 congregations by integrating multiple geospatial, knowledge-base, directory, and manually reviewed sources. Of these, 2,657 had active websites that were successfully crawled, producing a corpus of 154,454 webpages. We compared three approaches for detecting environmental actions: keyword retrieval followed by large language model (LLM) classification, semantic vector retrieval followed by LLM classification, and direct LLM classification classification without preliminary retrieval. Agreement with an expert human reviewer was lowest for keyword retrieval ($κ$ = 0.26), higher for semantic vector retrieval ($κ$ = 0.42), and similar for direct LLM classification ($κ$ = 0.40). Although semantic retrieval achieved the highest agreement, its retrieval recall was 0.87, indicating loss of relevant content before classification. Applied to the complete corpus, direct LLM classification identified at least one environmental action at 1,398 congregations (53%), providing greater coverage than either retrieval-based approach. These results demonstrate that preliminary retrieval can reduce computational cost but may exclude relevant information before it reaches the classifier. The framework provides a reproducible approach for extracting organization-level environmental information from unstructured web content that can be adapted to other institutions.
Comments22 pages, 6 figures, appendices