发表机构
Washington University in St. Louis School of Law(圣路易斯华盛顿大学法学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究旨在通过增强方法从公司组织文档提取数据,介绍DECODEM基准数据集,评估多种大语言模型提取管道,结果显示自动提取可行且性能因变量而异,标准化基准和评估证明前沿模型能高精度提取信息,凸显自动特征提取作用。
AI 中文摘要
许多实证法律研究依赖于将非结构化文本转化为结构化变量。在公司治理研究中,传统上依靠人工对章程等文件进行编码,成本高、难扩展且不透明。本文介绍了DECODEM,这是一组用于评估从组织文档中自动提取公司治理变量的基准数据集。利用这些数据集评估了几种在提示设计、任务分解和文档处理方面不同的大语言模型提取管道。结果表明,自动提取对许多条款来说在高精度下是可行的,性能因变量而异,更精细的提示策略和级联管道在某些情况下缩小了前沿模型和效率导向模型之间的差距。通过提供标准化基准和系统评估提取方法,证明当前前沿模型能从复杂公司文档中高精度提取有法律意义的信息,并表明自动特征提取在构建公司治理数据集方面有重要未来作用。
英文摘要
Much empirical legal research depends on translating unstructured text into structured variables. In corporate governance research as elsewhere, this translation has traditionally relied on human coding of documents such as charters and bylaws, a process that is costly, difficult to scale, and often opaque. This paper introduces DECODEM, a set of benchmark datasets for evaluating the automated extraction of corporate governance variables from organizational documents. The benchmarks pair randomly sampled corporate charters and bylaws with high-quality human annotations covering a range of governance provisions commonly studied in empirical work. Using these datasets, the paper evaluates several large-language-model extraction pipelines that vary in prompt design, task decomposition, and document handling. The underlying task consists of a set of document-level binary classification problems, one for each governance variable. The results show that automated extraction is feasible at a high level of accuracy for many provisions, with median performance near the upper bound across approaches. At the same time, performance varies systematically across variables, with a small number of provisions accounting for most of the remaining errors. More elaborate prompting strategies and cascading pipelines do not consistently improve performance for frontier models, but substantially narrow the gap between frontier and efficiency-oriented models in some settings, suggesting that pipeline design can partly substitute for model capability. By providing a standardized benchmark and a systematic evaluation of extraction methods, the paper demonstrates that current frontier models can extract legally meaningful information from complex corporate documents with high accuracy and suggests an important future role for automated feature extraction in constructing corporate governance datasets.