从政策文件到结构化调查响应:评估用于政策监测的大型语言模型
From Policy Documents to Structured Survey Responses: Evaluating Large Language Models for Policy Monitoring
浏览论文内容
中文总结 AI 辅助
本研究提出利用大型语言模型作为“AI受访者”从政策文档生成结构化调查响应,通过长上下文学习管道实现高效政策监测,实验显示结构化指标一致性达84-95%,凸显人机混合工作流的潜力。
中文摘要 AI 辅助
科学、技术和创新政策对竞争力至关重要,然而其多样性和规模使得对其进行一致性的映射和监测变得困难。现有方法严重依赖人工调查工作,这些工作成本高昂且难以在各国间扩展。大型语言模型(LLMs)为从冗长且非结构化的政策文件中提取和结构化信息提供了新的可能性。本文展示了将LLMs作为“AI受访者”用于从政策文本生成结构化调查响应的应用。我们开发了一个基于长上下文上下文学习的数据提取管道,将来自公共网络来源的信息映射到预定义的调查类别中,包括政策工具、目标群体和主题领域。该管道整合了一个验证步骤,使用辅助LLM评估相关性和证据,并与人工提供的响应进行比较。利用一个多国数据集,我们通过重叠度量和交叉验证评估了LLM生成输出与人工生成输出之间的一致性。结果表明,LLMs在结构化指标上实现了高度一致性(84-95%),而在自由文本字段中仍存在差异,模型往往提供更详细的过程性描述。这些发现凸显了人机混合工作流在政策监测中的潜力,既提高了效率和可扩展性,又保持了人工验证和情境化解释的必要性。
英文摘要
Science, technology, and innovation policies are crucial for competitiveness, yet their diversity and scale make them difficult to map and monitor consistently. Existing approaches rely heavily on manual survey efforts, which are costly and challenging to scale across countries. Large language models (LLMs) enable new possibilities for extracting and structuring information from long and unstructured policy documents. This paper presents an application of LLMs as "AI respondents" for generating structured survey responses from policy texts. We develop a data extraction pipeline based on long-context in-context learning to map information from public web sources into predefined survey categories, including policy instruments, target groups, and thematic areas. The pipeline integrates a validation step using a secondary LLM to assess relevance and evidence, alongside comparisons with human-provided responses. Using a multi-country dataset, we evaluate the alignment between LLM-generated and human-generated outputs through overlap measures and cross-validation. Results show that LLMs achieve high agreement for structured indicators (84-95%), while differences remain in free-text fields, where models tend to provide more detailed procedural descriptions. These findings highlight the potential of hybrid human-AI workflows for policy monitoring, improving both efficiency and scalability while maintaining the need for human validation and contextual interpretation.
发表机构
- VTT Technical Research Centre of Finland Ltd.(芬兰国家技术研究中心)
机构由 AI 辅助整理,请以论文原文为准。