从提案到验证效果:Praxa,一个受证据约束的受治理AI智能体执行框架
From Proposal to Verified Effect: Praxa, an Evidence-Bound Harness for Governed AI Agent Execution
- Independent Researcher(独立研究者)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
Praxa框架通过确定性准入、代理执行和外部回读等机制显式追踪智能体行动的授权到效果转换,提供四条证据路径验证其可靠性,但当前证据未证明性能优越性或生产效益。
AI中文摘要:
大语言模型智能体可以提出并执行行动,但提案、授权、调度、验证的外部效果以及服务推广是不同的声明。我们提出了Praxa,一个智能体框架,通过确定性准入、代理执行、外部回读、对账和审查推广,显式地表示这些状态。我们报告了四条证据路径。首先,作者运行的仓库本地审计在固定修订版上通过了1,027/1,027个单元测试和89/89个Workerd测试,覆盖了所有363个预期源文件,并达到了四个覆盖率下限;原始逐测试记录和独立复现不可用。其次,在提供商支持的Terminal-Bench Core 0.1.1试点中,跨越12个精选任务,基线和可靠性层各通过了17/36个严格试验。可靠性层使用了多37.49%的输入和多50.73%的输出令牌,因此该试点不支持优越性。第三,在调试后的两阶协调代理开发比较中,基线和源作者候选各完成了180/180个试验,具有相同的测量精度、完全封闭的崩溃恢复和零受保护违规。候选使用了少37.11%的令牌、低33.84%的估计端点成本和少11.63%的步骤;这并未确立改进的质量、延迟或生产行为。第四,部署的源/配置证据显示了有界的反思、召回核算、内存编译和工具健康路径,但没有生产结果提升。Praxa的支持贡献是一个受证据约束的架构,使权威到效果的转换显式且可测试。当前证据并未确立对抗性安全、生产安全、通用专家优越性、自主递归优化或用户利益。
英文摘要:
Large-language-model agents can propose and execute actions, but proposal, authority, dispatch, verified external effect, and serving promotion are different claims. We present Praxa, an agent harness that represents these states explicitly through deterministic admission, brokered execution, external read-back, reconciliation, and reviewed promotion. We report four evidence lanes. First, an author-run repository-local audit at a pinned revision passed 1,027/1,027 unit tests and 89/89 Workerd tests, instrumented all 363 expected source files, and met four coverage floors; raw per-test transcripts and independent reproduction are unavailable. Second, in a provider-backed Terminal-Bench Core 0.1.1 pilot across 12 curated tasks, baseline and reliability-layer arms each passed 17/36 strict trials. The reliability layer used 37.49% more input and 50.73% more output tokens, so the pilot does not support superiority. Third, in a post-debug, two-order coordination-proxy development comparison, baseline and a source-authored candidate each completed 180/180 trials with equal measured accuracy, full hermetic crash recovery, and zero protected violations. The candidate used 37.11% fewer tokens, 33.84% lower estimated endpoint cost, and 11.63% fewer steps; this does not establish improved quality, latency, or production behavior. Fourth, deployed source/configuration evidence shows bounded reflection, recall accounting, memory compilation, and tool-health paths, but no production outcome lift. Praxa's supported contribution is an evidence-bound architecture that makes authority-to-effect transitions explicit and testable. Current evidence does not establish adversarial security, production safety, general specialist superiority, autonomous recursive optimization, or user benefit.