发表机构
Joule(焦耳)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
AID框架通过统一物理、计算、网络和服务过程,定义了AI推理基础设施的学习问题,并提供了预测误差下界与状态约简条件,用于支持策略评估和干预响应。
AI 中文摘要
一个有用的AI推理基础设施模型必须指定系统状态、观察者可用的信息以及模型旨在支持的决策。我们引入AID(AI基础设施动态),这是一个用于描述跨物理、计算、网络和服务过程的耦合学习问题的框架。该公式允许结构化和可变大小的状态、异步观察、多个物理时间尺度以及对服务做出响应的需求。我们区分了在现有策略下支持预测的表示与在改变行动下保留服务结果的表示,并将两者与识别干预响应区分开来。两个分析结果描述了当可用观察无法区分模型时预测误差的下界,以及精确受控状态约简的充分条件。这些结果将既定的信息和状态抽象原则应用于AI基础设施。然后,我们描述了一个用于缓存表示、工作负载历史、测量可用性和施加行动的验证协议。
英文摘要
A useful model of AI inference infrastructure must specify the system state, the information available to an observer, and the decisions the model is intended to support. We introduce AID (AI Infrastructure Dynamics), a framework for describing this learning problem across coupled physical, computational, networking, and serving processes. The formulation allows structured and variable-size state, asynchronous observations, multiple physical timescales, and demand that responds to service. We distinguish representations that support prediction under an existing policy from those that preserve service outcomes under changed actions, and separate both from identifying intervention responses. Two analytical results describe a lower bound on prediction error when available observations cannot distinguish models and a sufficient condition for exact controlled state reduction. These results apply established information and state-abstraction principles to AI infrastructure. We then describe a validation protocol for cache representations, workload histories, measurement availability, and imposed actions.
Comments14 pages, 3 figures