arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.32184cs.AIcs.SYeess.SY

AI 驾驭:基础模型智能体在提案条件信息下的认证

AI Harness: Certification under Proposal-Conditioned Information for Foundation-Model Agents

Hailin Zhong, Shengxin Zhu

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出AI驾驭框架,证明基础模型智能体中提案条件信息对认证的必要性,通过鲁棒接口分析揭示压缩模型可认证性损失的条件,并扩展至时间域验证。核心贡献是表征模型-工具边界相关性何时不可或缺。

中文摘要 AI 辅助

基础模型智能体通常被建模为基于观测状态的政策。然而,在部署系统中,运行时可能仅在模型发出语义提案后才进行干预,这使得提案既是行动候选,又是由历史条件过程生成的决策时观测。我们表明,将这种结构压缩为仅状态提案包络可以在保留提案覆盖范围的同时破坏可认证性。在有限鲁棒接口中,压缩模型的可行核包含在历史增强核的物理投影中,且当每个提案条件压缩纤维保留共同的鲁棒安全干预时,压缩是无损的。即使提案和历史字母表大小恒定,这种差距也可能达到最大。相同的共同行动条件产生对偶结果:当当前提案分离需要不兼容干预的潜在模式时,观测当前提案可以恢复鲁棒可行性。我们使用精确有限信念和标准安全与可达性不动点将这些一步结果随时间扩展,将无限期操作可行性与有限最坏情况验证进展区分开来。受控模型在环测试在遥测或效果验证被移除或干预权限受限时重现了预测的障碍。因此,我们的贡献不是新的不动点演算,而是表征模型-工具边界处提案-历史相关性何时对认证是必要的。

英文摘要

Foundation-model agents are often modeled as policies over an observed state. In deployed systems, however, a runtime may intervene only after the model has emitted a semantic proposal, making the proposal both an action candidate and a decision-time observation generated by a history-conditioned process. We show that collapsing this structure into a state-only proposal envelope can preserve proposal coverage while destroying certifiability. In a finite robust interface, the viability kernel of the collapsed model is contained in the physical projection of the history-augmented kernel, and the collapse is lossless exactly when every proposal-conditioned collapsed fiber retains a common robust-safe intervention. This gap can be maximal even with constant-size proposal and history alphabets. The same common-action condition yields a dual result: observing the current proposal can restore robust feasibility when it separates latent modes requiring incompatible interventions. We extend these one-step results over time using exact finite beliefs and standard safety and reachability fixed points, separating indefinite operational viability from finite worst-case verified progress. Controlled model-in-the-loop tests reproduce the predicted obstructions when telemetry or effect verification is removed or intervention authority is restricted. Thus, our contribution is not a new fixed-point calculus, but a characterization of when proposal--history correlation at the model--tool boundary is necessary for certification.

↑