arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.30057cs.DCcs.OScs.PF

KREX:通过区域粒度独占性在共享GPU上进行并发内核基准测试

KREX: Concurrent Kernel Benchmarking on Shared GPUs via Region-Granular Exclusivity

Tianyu Feng, Haoxuan Yu, Tianyuan Wu, Lingyun Yang, Daocheng Ying, Yuxiao Wang, Ruibo Fan, Yinghao Yu, Guodong Yang, Liping Zhang, Wei Wang

首次发表
浏览论文内容

中文总结 AI 辅助

KREX通过区域粒度独占性实现共享GPU上的并发内核基准测试,在保持测量保真度的同时,将吞吐量提升至3.4倍,且时序膨胀极小。

中文摘要 AI 辅助

LLM智能体通过反复组合候选内核并在真实GPU上测量其运行时长来自动化GPU内核优化。现有系统通过为整个智能体会话或基准测试命令保留GPU来保证测量保真度。然而,这导致利用率低下,因为只有一小部分命令执行需要独占GPU访问。共享GPU可以恢复这种空闲容量,但会引入争用,从而损害测量保真度并误导智能体的搜索。我们提出KREX,一个支持区域粒度独占性的并发内核智能体基准测试运行时。KREX允许智能体在基准测试命令中标记涉及时序敏感操作的关键区域。运行时在标记区域内强制独占性,并允许在区域外并发执行,从而在保持测量保真度的同时实现高吞吐量。为了强制区域内独占性,KREX阻止新的竞争GPU提交,并在冻结兄弟进程和隔离CPU核心之前排空未完成的工作,从而保护GPU执行和驱动测量的主机线程。为了最大化区域外并发性,KREX在持久上下文进程中重用GPU上下文,以避免重复的、节点级串行的上下文创建。我们在NVIDIA和AMD GPU上评估了KREX。与命令粒度独占性基线相比,KREX提供了高达3.4倍的基准测试吞吐量,同时对于时长超过10毫秒、1毫秒和0.1毫秒的内核,p95时序膨胀分别仅为0.30%、1.58%和3.90%,可忽略不计。

英文摘要

LLM agents automate GPU kernel optimization by repeatedly composing candidates and measuring their duration on real GPUs. Existing systems preserve measurement fidelity by reserving a GPU for an entire agent session or benchmarking command. However, this results in poor utilization because only a small fraction of command execution requires exclusive GPU access. Sharing GPUs could recover this idle capacity, but introduces contention that compromises measurement fidelity and misdirects the agent's search. We present KREX, a runtime for concurrent kernel agent benchmarking with region-granular exclusivity. KREX lets agents mark critical regions involving timing-sensitive operations within a benchmarking command. The runtime then enforces exclusivity within marked regions and allows concurrent execution outside them, achieving high throughput while preserving measurement fidelity. To enforce in-region exclusivity, KREX blocks new competing GPU submissions and drains outstanding work before freezing sibling processes and isolating CPU cores, protecting both GPU execution and the host threads that drive measurements. To maximize off-region concurrency, KREX reuses GPU contexts in persistent context processes to avoid repeated, node-wide serialized context creation. We evaluate KREX on NVIDIA and AMD GPUs. Compared with command-granular exclusivity baselines, KREX delivers up to $3.4\times$ the benchmarking throughput with a negligible p95 timing inflation of $0.30\%$, $1.58\%$, and $3.90\%$ for kernels longer than 10 ms, 1 ms, and 0.1 ms, respectively.

发表机构

  • HKUST(香港科技大学)
  • Alibaba Group(阿里巴巴集团)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑