arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

需求约束的验证调试:一个冻结的四十亿参数本地模型作为候选生成器,置于具有验证与发布权限的外部验收层之下

Requirement-Bound Verified Commissioning: A Frozen Four-Billion-Parameter Local Model as a Candidate Generator under an External Acceptance Layer with Verification and Release Authority

Mehmet Iscan

arXiv 2609.30219首次发表:更新:

发表机构

Yıldız Technical University(耶尔德兹技术大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出一种机电调试验收协议,将候选生成与发布权限分离,用冻结的四十亿参数本地模型处理非确定性需求,并通过外部密封语法门控验证,实验表明能有效拒绝伪造计划,但存在少量错误发布风险。

AI 中文摘要

针对机电调试中的传感器坐标和极性绑定,开发了一种验收协议。候选生成与发布权限相分离。对于确定性解析器不支持的需求,路由至一个冻结的、具有四十亿参数的本地语言模型。仅当外部门控在密封语法下能够推导出两个事实时,计划才被发布。当符合条件时,向黄金标准用户请求一个规范答案。该协议在基准构建前固定的标准下,对144个任务进行了单次评估,这些任务由隔离的智能体上下文编写,且无法访问门控、语法或实验计划。确立了三个贡献。首先,候选生成和发布决策被分别测量。在22个被路由的不可回答任务中,有21个提交了伪造的就绪计划,且全部被拒绝。相同的83次发布在没有模型调用的情况下被复现。其次,在83次发布中未观察到错误发布。在独立同分布假设下,作为诊断,获得了单侧95% Clopper-Pearson上界0.0354,低于密封的5%阈值。然而,在基准之外的种子0下,随后在146次发布中记录了一次错误发布。第三,对错误用户答案的防护进行了表征。在96个可回答任务中,有13个任务的两个事实均从原始文本中绑定。在其余任务的431次配对中,有169次发布了错误答案,包括涉及坐标排除的失败。由于资格是根据答案密钥确定的,因此未测试可部署的提问策略。门控敏感性和真实用户行为未被测量。

英文摘要

An acceptance protocol is developed for sensor-coordinate and polarity binding in mechatronic commissioning. Candidate generation is separated from release authority. Requirements unsupported by a deterministic parser are routed to a frozen local language model with four billion parameters. Plans are released only when both facts can be derived by an external gate under a sealed grammar. One canonical answer is requested from a gold-standard user when eligible. The protocol was evaluated once under a criterion fixed before benchmark construction, on 144 tasks written by isolated agent contexts without access to the gate, grammar, or experimental plan. Three contributions are established. First, candidate generation and release decisions were measured separately. Fabricated ready plans were committed on 21 of 22 routed unanswerable tasks, and all were rejected. The same 83 releases were reproduced without model calls. Second, no false release was observed among 83 releases. A one-sided 95% Clopper-Pearson upper bound of 0.0354 was obtained as a diagnostic under an independent-and-identically-distributed assumption, below the sealed 5% threshold. However, one false release was subsequently recorded among 146 releases outside the benchmark at seed 0. Third, protection against incorrect user answers was characterized. Both facts were bound from the original text on 13 of 96 answerable tasks. Incorrect answers were released in 169 of 431 pairings on the remaining tasks, including failures involving coordinate exclusion. A deployable questioning policy was not tested because eligibility was determined from the answer key. Gate sensitivity and real user behavior were not measured.

Comments42 pages, 6 figures, 15 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑