AI 中文总结
SCALE-Sim EVA是一款面向IR感知加速器建模的可扩展、可可视化、可适配的模拟框架,通过张量命令与命令分解机制计算周期时序,可生成轨迹用于分析执行相关指标,解决了固定加速器模拟器难以扩展的问题。
AI 中文摘要
现代AI加速器日益结合异构计算单元、分层存储器、本地缓冲区及专用数据传输路径,这种多样性使得固定加速器模拟器难以扩展至其原始执行模型之外。我们提出SCALE-Sim EVA,这是一个面向IR感知加速器建模的可扩展、可可视化且可适配的模拟框架。EVA将工作负载表示为基于张量的命令,跟踪运行时张量的放置与就绪状态,并在可组合的硬件组件上执行命令,这些硬件组件带有用户定义的功能单元和内存行为。EVA并非重放地址级周期轨迹,而是根据张量就绪状态、硬件单元可用性以及建模的操作或传输延迟来计算周期时序。其命令分解机制弥合了编译器级IR粒度与硬件级执行粒度,允许全局张量操作在硬件层次结构中被下推并在本地调度。EVA还会生成开放的命令和存储轨迹,用于可视化分析执行时间线、内存占用及资源争用情况。
英文摘要
Modern AI accelerators increasingly combine heterogeneous compute units, hierarchical memories, local buffers, and specialized data movement paths. This diversity makes fixed accelerator simulators difficult to extend beyond their original execution model. We present SCALE-Sim EVA, an extensible, visualizable, and adaptable simulation framework for IR-aware accelerator modeling. EVA represents workloads as tensor-based commands, tracks runtime tensor placement and readiness, and executes commands on composable hardware components with user-defined functional units and memory behavior. Instead of replaying address-level cycle traces, EVA computes cycle timing from tensor readiness, hardware-unit availability, and modeled operation or transfer latency. Its command decomposition mechanism bridges compiler-level IR granularity and hardware-level execution granularity, allowing global tensor operations to be lowered and scheduled locally across a hardware hierarchy. EVA also emits open command and storage traces for visual analysis of execution timelines, memory occupancy, and resource contention.