arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ARCHER:用于可执行法规的智能规则与合规框架

ARCHER: Agentic Rule and Compliance Harness for Executable Regulations

Chiraag Singh Anand, Xue Wen Tan, Lionel Teo, Eric Tan

arXiv 2607.25566首次发表:更新:

发表机构

Infocomm Media Development Authority of Singapore(新加坡信息通信媒体发展局)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对建筑合规验证难题,提出ARCHER这一测试驱动、确定性编排的多智能体程序合成框架,能从法规代码生成验证代码,经评估其在多骨干模型上准确率高,成本-准确率分析显示可降低成本实现数据主权合规检查。

AI 中文摘要

验证建筑合规性需要对照大型建筑信息模型(BIM)设计验证数千条规则,既费力又成本高且不可扩展。现有自动合规检查器(ACC)难以在不同场景中通用,且许多是专有的。我们引入ARCHER,一个测试驱动、确定性编排的多智能体程序合成框架,能从法规实践代码生成可审计的验证代码,实现透明、可适应且可扩展的合规检查。通过在新数据集上评估六种智能体复杂度递增的框架,ARCHER的确定性多智能体编排对每个骨干模型都实现最高准确率,成本-准确率分析表明自托管开放权重模型使用ARCHER框架能以四分之一成本达到前沿API准确率的97.8%,使数据主权合规检查可行。

英文摘要

Verifying building compliance requires validating thousands of rules against large Building Information Modeling (BIM) designs, which is laborious, capital-intensive, and unscalable. Existing Automated Compliance Checkers (ACCs) are often difficult to generalize across different scenarios, as they are typically developed for highly specific rule sets and use cases. In addition, many ACCs are proprietary, meaning the underlying verification code is not released to end users, so users cannot verify whether their regulatory intent can be accurately captured. We introduce ARCHER (Agentic Rule and Compliance Harness for Executable Regulations), a test-driven, deterministically orchestrated multi-agent program-synthesis harness that generates auditable verification code from regulatory Codes of Practice, enabling transparent, adaptable, and scalable compliance checking. To characterize what makes agentic synthesis work, we evaluate a taxonomy of six harnesses of increasing agentic sophistication across four backbone models, spanning realistic data-governance tiers (from frontier third-party APIs to a fully on-premise open-weights model) on a novel dataset derived from real-world compliance scenarios. ARCHER's deterministic multi-agent orchestration achieves the highest accuracy for every backbone, improving mean union accuracy by 82% over a naive single-pass prompting baseline. Our cost-accuracy analysis further shows that using the ARCHER harness, a self-hosted open-weights model can reach 97.8% of frontier-API accuracy at a quarter of the cost, making data-sovereign compliance checking practical.

CommentsAccepted at NeurIPS 2026 Workshop on AI for Verifiable Coding

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑