发表机构
Daimler Truck AG; Technical University of Munich(戴姆勒卡车股份公司; 慕尼黑工业大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
RunSoC 2.0是异构MPSoC中汽车软件任务调度与分配的可定制框架,支持多种求解后端,可生成最优调度、暴露架构瓶颈,CP-SAT在硬实时调度中表现更优
AI 中文摘要
集中式汽车架构正日益将计算密集型工作负载整合到异构多处理器片上系统(MPSoC)中,这带来了严格的执行、内存和通信约束。本文提出RunSoC 2.0,这是一个用于异构MPSoC上任务调度与分配的早期设计空间探索的可定制框架。RunSoC 2.0以针对同构硬件分配的RunSoC 1.0为基础,通过对处理器特定执行时间、集群级组织和领域特定处理属性进行建模,将框架扩展到异构平台。它将任务集表示为受严格端到端延迟和核心亲和性约束的有向无环图(DAG),并将任务调度与分配表述为多目标优化问题,以最小化分层内存预算违反和核心间/集群间通信惩罚。该框架支持多种求解后端,包括COIN-OR Branch and Cut(CBC)、Google OR-Tools CP-SAT和遗传算法(GA),可实现精确方法、约束编程和元启发式方法的比较评估。我们使用10至500个任务的合成汽车任务集对RunSoC 2.0进行评估,这些任务集被映射到代表性异构MPSoC,包括瑞萨R-Car V4H、英伟达Jetson AGX Orin和德州仪器TDA4VM。结果表明,RunSoC 2.0可生成可行且最优的调度,暴露架构瓶颈,并支持平台替代方案的快速比较。值得注意的是,在受严格约束的硬实时调度实例中,CP-SAT始终优于CBC和GA。通过纳入集群感知通信和内存建模,RunSoC 2.0提高了早期MPSoC分析的真实性,同时保留了大型汽车工作负载的实用求解时间。
英文摘要
Centralized automotive architectures increasingly consolidate compute-intensive workloads onto heterogeneous Multi-Processor System-on-Chip (MPSoC), creating strict execution, memory, and communication constraints. This paper presents RunSoC 2.0, a customizable framework for early-stage design-space exploration of task scheduling and allocation on heterogeneous MPSoCs. Building on RunSoC 1.0, which targeted allocation on homogeneous hardware, RunSoC 2.0 extends the framework to heterogeneous platforms by modeling processor-specific execution times, cluster-level organization, and domain-specific processing properties. It represents task sets as directed acyclic graphs (DAGs) subjected to strict end-to-end latency and core-affinity constraints, and formulates task scheduling and allocation as a multi-objective optimization problem that minimizes hierarchical memory-budget violations and inter-core/inter-cluster communication penalties. The framework supports multiple solving backends, including COIN-OR Branch and Cut (CBC), Google OR-Tools CP-SAT, and a Genetic Algorithm (GA), enabling comparative evaluation of exact, constraint-programming, and meta-heuristic approaches. We evaluate RunSoC 2.0 using synthetic automotive task sets ranging from 10 to 500 tasks, mapped to representative heterogeneous MPSoCs, including the Renesas R-Car V4H, NVIDIA Jetson AGX Orin, and TI TDA4VM. The results show that RunSoC 2.0 can generate feasible and optimal schedules, expose architectural bottlenecks, and support rapid comparison of platform alternatives. Notably, CP-SAT consistently outperforms both CBC and the GA across tightly constrained hard real-time scheduling instances. By incorporating cluster-aware communication and memory modeling, RunSoC 2.0 improves the realism of early-stage MPSoC analysis while retaining practical solution times for large automotive workloads. (..)
CommentsPreprint for CASA@ECSA2026 workshop track