arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.21503cs.SEcs.CRcs.LG

BeTaL-GBI:几何置信接口的准入感知基准调优与全栈验证

BeTaL-GBI: Admission-Aware Benchmark Tuning and Full-Stack Verification of Geometric Belief Interfaces

Alvin Spivey, Yu Huang

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出 BeTaL-GBI 等改进方案,验证几何置信接口的架构断言,优化基准调优与见证机制,提升验证底层的可信度与审计能力。

中文摘要 AI 辅助

当验证底层能暴露自身断言中的错误而非仅模型输出时,其可信度更高。GBI-DCSE v3 证伪了一项架构断言:所报告的 Fisher 值 ε ~ 0.066 仅在切片 [ε, 3, 4, 5] 上满足 κ² ≤ 10⁴ 的预算,而完整框 [ε, 20]⁴ 则要求 ε ~ 0.326472。本勘误强调企业验证架构能否在保持断言可审计性的同时,隔离接口故障、任务能力、政策准入性与控制完整性。BoundaryBench v0.1 建立了基准:Qwen3-4B-Instruct-2507 完成了 768 次冻结执行,但 0% 通过了合约(369 次解析失败,399 次验证失败),限制了下游选择性指标。本伴随研究评估三项连续改进:其一,BeTaL-GBI v0.2 采用 LLM 闭环的基准调优,针对 2218750380 个网格点,将格式准入与条件性能分离(ρ_adm = N_admitted/N;ρ_task = N_verified/N_admitted);经模式修复后,无模型反馈搜索实现了 2.87% 的保留目标均值差距,优于无反馈基线(13.61%、11.46%)。其二,GBI v2 将静态密钥替换为与参考无关的见证状态 W 和政策 P;在 512 个合成任务中,16 门政策检测到全部 116 个注入的严重矛盾并接受全部 99 条干净记录(宽分母:4.27%),且零静默提升阻止了幻觉器与证据伪造代理。其三,GBI-DCSE v3 将 99 项断言映射为机器可读证据:96 项可测试断言中的 95 项通过,148 项独立检查无失败;该工具在 62 种配置下测试了签名账本、PBFT 法定人数与飞地伪造。在合成条件下,GBI-DCSE 是一种选择性、政策版本化、自审计的测试与路由底层。

英文摘要

A verification substrate is more credible when exposing errors in its own claims, not just model outputs. GBI-DCSE v3 falsified an architectural claim: the reported Fisher value epsilon ~ 0.066 satisfies the kappa^2 <= 10^4 budget only on the slice [epsilon, 3, 4, 5], while the full box [epsilon, 20]^4 requires epsilon ~ 0.326472. This erratum highlights whether an enterprise verification architecture can isolate interface failure, task competence, policy admissibility, and control integrity while keeping claims auditable. BoundaryBench v0.1 established the baseline: Qwen3-4B-Instruct-2507 completed 768 frozen executions, but 0% cleared the contract (369 failed parsing, 399 failed validation), limiting downstream selectivity metrics. This companion study evaluates three successive improvements. First, BeTaL-GBI v0.2 applies Benchmark Tuning with an LLM-in-the-loop over 2,218,750,380 grid points, separating format admission from conditional performance (rho_adm = N_admitted/N; rho_task = N_verified/N_admitted). Following schema repair, a model-free feedback search achieves a 2.87% mean held-out target gap, outperforming non-feedback baselines (13.61%, 11.46%). Second, GBI v2 swaps static keys for a reference-independent witness state W and policy P. Across 512 synthetic tasks, a 16-gate policy detects all 116 injected severe contradictions and accepts all 99 clean records (broad denominator: 4.27%). Hallucinator and evidence-forger surrogates are blocked with zero silent promotions. Third, GBI-DCSE v3 maps 99 claims to machine-readable evidence: 95 of 96 testable claims pass, with 148 standalone checks executed without failure. The harness exercises signed ledgers, PBFT quorums, and enclave forgery across 62 configurations. Under synthetic conditions, GBI-DCSE is a selective, policy-versioned, self-auditing test and routing substrate.

发表机构

  • Light Imaging Technologies, Inc.(光成像技术公司)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑