零指令传感器读取:五阶段软处理器上的寄存器映射外设与硬件PWM
Zero-Instruction Sensor Reads: Register-Mapped Peripherals and Hardware PWM on a Five-Stage Soft Processor
中文总结 AI 辅助
本文针对五阶段软处理器,以反应轮自平衡自行车控制回路为对象,通过寄存器映射外设、硬件PWM卸载实现零指令传感器读取,降低指令数、简化软件,两种配置的控制循环周期大幅小于20ms驱动帧。
中文摘要 AI 辅助
本文针对一款五阶段软处理器开展了面向应用驱动的专用化案例研究,以反应轮自平衡自行车的内部控制回路为评估对象。研究从一款MIPS架构风格的定制32位RISC核心出发,通过两种方式对设计进行专用化:其一,将两个高频访问的外设输入直接映射到架构寄存器状态,由硬件每个周期写入,且通过寄存器文件的写端口结构独占访问,无需仲裁;其二,将四个周期PWM通道卸载到硬件,由四个导出寄存器持续驱动,完全移除了软件中的周期驱动任务。由于外设值可作为普通寄存器操作数寻址,控制回路中的全部10次传感器读取无需专用指令和专用周期,可融入已执行的算术运算;而内存映射等效方案则需要为每次快照执行显式加载,额外消耗5条指令和周期。驱动路径同样将波形维护完全从软件中移除。由于这些扩展及与其一同部署的单周期阵列乘法器未在单个归档构建中同时存在,因此报告两种配置:一种为归档配置,其最坏情况循环为91周期;另一种为匹配部署系统的集成配置,为43周期。相对于20ms的驱动帧,这两种配置的余量分别约为7300倍和15000倍。两种情况的余量均远宽于满足实时性要求所需,因此该专用化并非实时合规的必要条件;其价值在于减少指令数量和简化软件,而非确定性,确定性已由片上单周期I/O区域提供。零指令传感器读取与该选择无关:乘法器无法影响外设读取是否需要专用指令。
英文摘要
We present a case study in application-driven specialization of a five-stage soft processor, evaluated on the inner control loop of a reaction-wheel self-balancing bicycle. Starting from a custom 32-bit RISC core in the MIPS tradition, we specialize the design in two ways. First, two frequently accessed peripheral inputs are mapped directly into architectural register state, written every cycle by hardware and owned exclusively through the register file's write-port structure rather than by arbitration. Second, four periodic PWM channels are offloaded to hardware and driven continuously from four exported registers, removing periodic actuation from software entirely. Because peripheral values are addressable as ordinary register operands, all ten sensor reads in the control loop cost no dedicated instruction and no dedicated cycle, folding into arithmetic that executes anyway; the memory-mapped equivalent requires an explicit load per snapshot and costs five extra instructions and cycles. The actuation path likewise removes waveform maintenance from software entirely. We report two configurations, because the extensions and the single-cycle array multiplier they were deployed alongside are not present together in a single archived build: an archived configuration, whose worst-case loop is 91 cycles, and the integrated configuration matching the deployed system, at 43 cycles. Against a 20 ms actuation frame these are margins of roughly 7,300x and 15,000x. The deadline is met by so wide a margin in either case that the specialization was not necessary for real-time compliance; its value lies in instruction count and software simplicity, not in determinism, which an on-chip single-cycle I/O region already provides. The zero-instruction sensor read is independent of that choice: the multiplier cannot affect whether a peripheral read needs an instruction of its own.