AI 中文总结
提出数字治理框架的任务替代框架,将关卡作为可执行合约,通过前置部署工程和DGF-Bench验证,证明智能体可替代指定治理审查任务的人工执行。
AI 中文摘要
企业治理需要决策、证据和可问责的权威;它并不要求每项审查任务都保留其当前的人工执行方式。我们为数字治理框架(DGF)开发了一个任务替代框架,将每个关卡视为一个可执行的合约。替代需要充足的可获取信息、有效的决策和权威检查,并且在计入异常、验证、纠正和维护之后,总人工工作量有所减少。我们推导出一个剩余工作量阈值,并说明了为什么自动化大多数案例仍可能增加劳动量。前置部署工程将这些条件与一个由智能体、规则引擎、证据服务和升级机制组成的架构联系起来。DGF-Bench 提供了来自 300 个合成项目和 899 次可评估的模型-项目运行的可控证据。Gemini 3.8 Flash、GPT-5.6 Luna 和 DeepSeek v4.1 Flash 的严格关卡成功率分别为 94.98%、83.29% 和 74.18%;完整路径成功率分别为 76.92%、42.33% 和 24.67%。一个确定性控制在给定规则和结构化事实的情况下通过了全部 1,700 个关卡,从而将比较定位在给定决策内核的执行上。证据审计和 135 次重复运行区分了正确决策与可靠执行。一个文档反例确立了信息充分性障碍。这些结果支持了用智能体和软件替代指定治理审查任务的人工执行在技术上的可行性。该框架基于固定产出和质量下所需的完整人工工作量,规定了一项劳动力测试;当前的测量涉及审查性能。来源、档案、轨迹和分析均为公开。
英文摘要
Enterprise governance requires decisions, evidence, and accountable authority; it does not require every review task to retain its current human implementation. We develop a task-substitution framework for Digital Governance Frameworks (DGF), treating each gate as an executable contract. Substitution requires sufficient accessible information, valid decision and authority checks, and a reduction in total human work after exceptions, verification, correction, and maintenance are counted. We derive a residual-work threshold and show why automating most cases can still increase labor. Forward deployed engineering connects these conditions to an architecture for agents, rule engines, evidence services, and escalation. DGF-Bench supplies controlled evidence from 300 synthetic projects and 899 evaluable model-project runs. Gemini 3.8 Flash, GPT-5.6 Luna, and DeepSeek v4.1 Flash achieve strict gate success of 94.98%, 83.29%, and 74.18%; complete-route success is 76.92%, 42.33%, and 24.67%. A deterministic control passes all 1,700 gates given the supplied rules and structured facts, locating the comparison in execution of a supplied decision kernel. Evidence audits and 135 repeated runs distinguish correct decisions from reliable execution. A document counterexample establishes an information-sufficiency obstruction. These results support the technical feasibility of replacing human execution of specified governance-review tasks with agents and software. The framework specifies a workforce test based on the complete human effort required at fixed output and quality; the present measurements concern review performance. Sources, dossiers, traces, and analyses are public.
Comments28 pages, 5 figures, 12 tables. Code, datasets, generated documents, and model traces available at https://github.com/jeremy1392/dgf-agentic-bench