发表机构
Peking University; State Key Laboratory of General Artificial Intelligence, BIGAI; University of Chinese Academy of Sciences(北京大学; 通用人工智能国家重点实验室、北京智源人工智能研究院; 中国科学院大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出GNRS-Search框架,将规范规则生成为接地规范规则合成问题,通过MCMC采样优化AOG,在两个基准上提升规则质量,为受监管环境的合规智能体提供基础范式。
AI 中文摘要
机构章程、职场政策等规范规则必须兼具人类可读性,且可依据实际环境记录进行操作验证。然而,当前的语言生成与结构化输出基准主要奖励表面流畅度或模式合规性,对操作接地性的测试极为薄弱,这形成了关键漏洞:标准语言模型会生成听起来合理的政策,但在执行时失败,因为它们依赖不可用的数据日志或未对齐的范围。为解决这一挑战,我们将该问题形式化为接地规范规则合成(Grounded Normative Rule Synthesis, GNRS),并引入GNRS-Search框架,该框架利用马尔可夫链蒙特卡洛(Markov Chain Monte Carlo, MCMC)采样优化离散的五槽与或图(And-Or Graph, AOG)。通过明确将中间操作结构与最终文本生成解耦,该方法将可执行性与写作风格分离,允许在表面实现前定位规则失败。我们在GNRS-Bench上评估了该方法,该基准涵盖8个场景族中的116个受控目标;以及RealCharter-Bench,用于评估向53个源自真实政策任务的迁移,这些任务带有隐藏的源条款。GNRS-Search将平均评分表质量从68.8%提升至81.0%,并在公开的可执行性综合指标下排名第一;系统的槽干预实验证实,性能提升源于稳健的操作逻辑,而非修辞调整。最终,通过将自动规则起草转化为可检查的搜索问题,本研究为在受监管环境中部署可验证、合规就绪的个人智能体提供了基础范式。
英文摘要
Normative rules like institutional charters and workplace policies must be both human-readable and operationally verifiable against actual environment records. However, current language generation and structured-output benchmarks primarily reward surface fluency or schema compliance, leaving operational grounding weakly tested. This creates a critical vulnerability where standard language models generate plausible-sounding policies that fail during enforcement because they rely on unavailable data logs or misaligned scopes. To address this challenge, we formalize the problem as Grounded Normative Rule Synthesis (GNRS) and introduce GNRS-Search, a framework that utilizes Markov Chain Monte Carlo (MCMC) sampling to optimize a discrete, five-slot And-Or Graph (AOG). By explicitly decoupling intermediate operational structure from final prose generation, this method isolates executable feasibility from writing style and allows rule failures to be localized prior to surface realization. We evaluate our approach on GNRS-Bench, a benchmark spanning 116 controlled goals across eight scene families, and RealCharter-Bench, which evaluates transfer to 53 real-derived policy tasks with hidden source clauses. GNRS-Search raises average rubric quality from 68.8% to 81.0% and ranks first under a disclosed executable composite metric, while systematic slot interventions confirm that performance gains stem from robust operational logic rather than rhetorical tuning. Ultimately, by transforming automated rule drafting into an inspectable search problem, this work provides a foundational paradigm for deploying verifiable and compliance-ready personal agents within regulated environments.