GPS-Bench:用于自动化政策分析的治理政策基准
GPS-Bench: A Governance Policy Benchmark for Automating Policy Analysis
- Lida Safety(丽达安全)
- ERA
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
GPS-Bench是基于公开证据的治理政策模拟基准,通过受控对比验证多主体模拟等方法对政策结果预测与解释的提升作用,为相关研究提供了统一实证环境。
AI中文摘要:
政策分析不仅需要预测提案是否会通过,还需要识别哪些主体会受到影响、这些主体如何回应以及后续会产生什么结果。基于大语言模型(LLM)的政策模拟可规模化地对这些过程进行建模,但当合理行为从未与观测结果进行比较时,其有效性难以确立。我们引入GPS-Bench,这是一个基于证据的治理政策模拟基准,它利用立法记录、游说披露、监管文件、公司备案、经济数据及其他公开证据,将政策与相关主体、主体行为及下游影响关联起来。主体是从时间记录中重建的,而非作为原型被提示,因此角色是具有来源的证据对象;人工标注的集合构成黄金评估集,而由独立LLM从检索证据中标记的案例被视为白银监督,永远不会作为测试标签。由于每种推理模式都读取相同的基础状态并输出相同的模式,GPS-Bench将“多主体模拟是否有帮助?”转化为一项受控比较:我们在一个政策状态上对比联合推理、独立且通信的主体智能体、基于图的方法以及权重级微调。对基础记录进行微调可获得最强的主体层面影响预测,而分解方法并未优于它;分解所增加的是机制。智能体持有私人、不相同的证据,每个智能体都能看到自身的披露条款,并以具体的联合提案、自身提供的内容、需要回报的条件以及共同行动为何优于单独行动的理由来与指定伙伴互动,因此形成的联盟可与记录中持有的承诺进行核对。GPS-Bench因此为研究证据、主体建模和多主体互动在何种情况下能改善政策结果的预测与解释提供了一个共同的实证环境。
英文摘要:
Policy analysis requires more than predicting whether a proposal will pass: it requires identifying who will be affected, how those actors respond, and what follows. LLM-based policy simulations model these processes at scale, but their validity is hard to establish when plausible behaviour is never compared with observed outcomes. We introduce GPS-Bench, an evidence-grounded benchmark for governance policy simulation that links policies to relevant actors, actor actions and downstream impacts using legislative records, lobbying disclosures, regulatory documents, corporate filings, economic data and other public evidence. Actors are reconstructed from the dated record rather than prompted as archetypes, so a persona is an evidence object with provenance; a human-annotated pool forms the Gold evaluation set, while cases labelled by a separate LLM from retrieved evidence are treated as Silver supervision and never as test labels. Because every inference mode reads the same grounded state and emits the same schema, GPS-Bench turns "does multi-agent simulation help?" into a controlled comparison: we contrast joint reasoning, independent and communicating actor agents, graph-based methods and weight-level fine-tuning over one policy state. Fine-tuning on the grounded record gives the strongest actor-level impact prediction, and decomposition does not beat it; what decomposition adds is mechanism. Agents hold private, non-identical evidence, each seeing its own exposure clause, and address named partners with concrete joint proposals, what they offer, what they need in return, and why acting together beats acting alone, so the coalitions that form can be checked against the commitments the record holds. GPS-Bench therefore gives a common empirical setting for studying when evidence, actor modelling and multi-agent interaction improve the prediction and interpretation of policy outcomes.