AI 中文总结
本文针对自适应生成式AI违背静态模型生命周期假设的治理问题,提出含离散与连续时间的遥测架构及对抗性模型隐藏的6种攻击对策,确立需结合连续遥测与定期审计的双机制架构,符合传统MRM要求。
AI 中文摘要
模型风险管理(MRM)指南假设模型生命周期是静态的,即模型被开发、独立验证和部署后不会再进行自主修改。而持续自适应的生成式AI系统(这类模型会在生产部署期间自主更新权重)从根本上违背了这一假设,使得时点验证不再适用。本文分两部分解决由此产生的治理问题:第一部分为这类模型开发了同时工作在离散时间和连续时间的严谨遥测架构,为审计目的建立了最小充分统计量,为离散权重序列构建了防篡改的默克尔链,通过伊藤公式推导了对应的连续时间泛化,并提出了基于KL散度停止时间的事件驱动日志记录,该方法兼具计算可行性和验证意义;第二部分探讨当模型提供者具有对抗性时会发生什么,部署此类模型的企业有强烈动机隐藏会触发强制验证审查的学习更新,我们将此正式定义为模型隐藏问题,并针对第一部分的架构系统分类了6种不同的攻击策略,涵盖离散和连续时间,为每种策略提供了正式的对策。两部分共同确立了双机制架构,其中连续遥测是必要但不充分的,它会缩小但永远无法取代定期侵入式审计。该框架与模型架构无关,旨在满足传统MRM的三大支柱要求。
英文摘要
Model risk management (MRM) guidance assumes a static model lifecycle, in which models are developed, independently validated, and implemented without further autonomous modification. Continually self-adapting generative AI systems --- models that update their own weights during production deployment --- fundamentally violate this assumption and render point-in-time validation inadequate. This paper addresses the resulting governance problem in two parts. Part I develops a rigorous telemetry architecture for such models, operating simultaneously in discrete and continuous time. We establish a Minimal Sufficient Statistic for audit purposes, construct a tamper-evident Merkle chain for discrete weight sequences, derive the appropriate continuous-time generalization via the Ito formula, and propose event-driven logging via KL divergence stopping times that is both computationally tractable and meaningful for validation. Part II asks what happens when the model provider is adversarial. A firm deploying such a model has strong incentives to conceal learning updates that would trigger mandatory validation review. We formalize this as the Model Hiding Problem and provide a systematic taxonomy of six distinct attack strategies against the Part I architecture, spanning discrete and continuous time, with a formal countermeasure for each. Together the two parts establish a dual-regime architecture in which continuous telemetry is necessary but not sufficient, narrowing---but never replacing---periodic invasive audit. The framework is model-architecture-agnostic and is designed to satisfy the three pillars of traditional MRM.