arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.17899cs.AR

SEAM-V:一种具有后端可见执行包上下文的混合解耦RISC-V向量处理器,用于持续向量吞吐量

SEAM-V: A Hybrid-Decoupled RISC-V Vector Processor with Backend-Visible Packet Semantics and Source-Lifetime-Aware Scheduling

Weiying Wang, Zhiwei Zhang

首次发表
浏览论文内容

中文总结 AI 辅助

针对深度学习和科学计算中对处理器性能的需求,提出SEAM-V这一混合解耦的RISC-V向量执行架构,通过任务级解耦等形成执行包流,实现同-EP候选危险抑制等,相比紧密耦合实现有显著加速,不同应用组加速效果各异。

中文摘要 AI 辅助

深度学习和科学计算中的数据并行工作负载持续推动对更高处理器吞吐量、能源效率和可扩展性的需求。RISC-V向量扩展(RVV)通过与向量长度无关的编程模型支持可扩展执行。然而,许多紧密耦合实现仍依赖标量核一次提供一条向量指令,导致执行易受向量指令供应间隙、标量侧推进延迟以及短向量、循环尾部和控制/内存交错阶段保守依赖处理的影响。本文提出了SEAM-V,一种用于RVV的混合解耦向量执行架构。SEAM-V通过任务级解耦、本地指令供应和VLIW风格打包形成连续的执行包(EP)流。一个EP被序列化为单个请求后,其EP标识和请求绑定的预取上下文对动态向量后端仍然可见,实现同-EP候选危险抑制和请求绑定预取。混合调度路径在满足所需安全条件时还可提供有限的跨EP向量重叠。跨EP依赖、未被EP契约豁免的依赖、资源冲突和内存排序仍由后端动态管理。精确到周期的RTL评估表明,与基于Ara的紧密耦合RVV实现(TC)相比,SEAM-V在17个代表性内核上实现了1.34倍的几何平均加速。一维可变AVL、BLAS和矩阵以及固定大小应用组分别实现了1.50倍、1.25倍和1.27倍的加速。在AVL = 32时,六个一维向量内核的几何平均加速接近3倍。

英文摘要

Data-parallel workloads in deep learning and scientific computing continue to increase the demands on processor throughput, energy efficiency, and scalability. The RISC-V Vector Extension (RVV) supports scalable execution through a vector-length-agnostic model, yet many tightly coupled implementations still rely on the scalar core to supply vector instructions individually and are therefore constrained by instruction supply, scalar-side progress, memory stalls, and conservative dependence management in short-vector, loop-tail, and control/memory-interleaved scenarios. This paper presents SEAM-V, a hybrid-decoupled RVV processor that uses task-level decoupling, local instruction supply, and VLIW-style packing to form a continuous stream of execute packets (EPs). During EP formation and request serialization, SEAM-V preserves the association between prefetch intent and the corresponding load to support request-bound prefetching, while lane-level source-read completion is used to release pure write-after-read (WAR) dependences early; other dependences remain governed by conventional mechanisms. Relative to an Ara-based tightly coupled baseline (TC), SEAM-V achieves a geometric mean speedup of 1.38x across 17 representative kernel configurations. The one-dimensional vector, BLAS and matrix, and fixed-size application workload groups achieve 1.56x, 1.35x, and 1.23x, respectively. Synthesis and power analysis show that SEAM-V increases total cell area by 4.29% and geometric mean runtime power by 17.30%, while reducing task energy by 12.70%, demonstrating improved sustained execution efficiency and task-level energy efficiency with limited area overhead.

补充信息

↑