arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.35508cs.OScs.AR

Argus:智能体驱动、参考校准、树引导的系统软件级瓶颈定位

Argus: Agentic, Reference-Calibrated, Tree-Guided, System-Software-Level Bottleneck Localization

Vlad-Petru Nitu, Harsh Songara, Konstantinos Sgouras, Spiros Galanopoulos, Konstantinos Kanellopoulos, Onur Mutlu

首次发表 更新
浏览论文内容

中文总结 AI 辅助

针对操作系统瓶颈定位依赖人工、现有工具易干扰操作的问题,本研究提出基于LLM的智能体分析工具Argus,通过参考校准和树状执行路径结构实现精准内核代码路径定位,错误率较基线低19倍,诊断耗时约31秒。

中文摘要 AI 辅助

操作系统(OS)代码可占据CPU执行时间的很大一部分。首先,随着应用逻辑被卸载到异构加速器(如GPU),CPU越来越多地充当协调者,将计算周期耗费在驱动调用、数据移动和同步上,而非应用代码中。其次,诸如无服务器函数之类的工作负载会频繁调用OS服务。与此同时,OS是一个跨越许多子系统(如内存管理、网络)的复杂代码库,使得定位导致性能下降的特定代码路径十分困难。现有分析工具提供的测量结果需要人工解读(例如perf和Intel VTune),或者在进行广泛插桩时会干扰短时长操作(例如ftrace)。因此,诊断OS瓶颈可能需要重复的内核插桩和人工解读。\n我们推出了Argus,这是一款基于大语言模型(LLM)的智能体分析工具,能够生成插桩代码并自主推理潜在的OS级瓶颈。Argus整合了两个关键机制:(i)一种校准方法,包括从空闲系统收集测量数据并将其用作参考点,以发现潜在瓶颈;(ii)一种基于树的数据结构,用于表示不同的OS执行路径,提升智能体的瓶颈定位准确率。Argus的目标是识别特定的内核代码路径,而非停留在子系统级别的诊断。在两个案例研究中,我们使用Argus自主发现了内存管理子系统中存在的瓶颈,这些瓶颈分别由(i)一个THP(透明巨页)干扰程序与其他应用共同运行,以及(ii)产生不同类型页错误的应用所导致。Argus产生的错误深层路径诊断数量比所评估的最强的、缺少参考校准的基于LLM的基线少19倍,同时保持了较低的诊断耗时(约31秒)。

英文摘要

Operating system (OS) code can account for a substantial share of CPU execution time. First, as application logic is offloaded to heterogeneous accelerators (e.g., GPUs), the CPU increasingly acts as an orchestrator, spending cycles in driver calls, data movement, and synchronization rather than in application code. Second, workloads such as serverless functions frequently invoke OS services. At the same time, the OS is a complex codebase spanning many subsystems (e.g., memory management, networking), making it hard to localize the specific code path responsible for a slowdown. Existing profilers expose measurements that require interpretation(e.g., perf and Intel VTune) or can perturb short operations when extensively instrumented (e.g., ftrace). Diagnosing OS bottlenecks can therefore require repeated kernel instrumentation and manual interpretation. We introduce Argus, an agentic LLM-based profiler that produces instrumentation code and autonomously reasons over potential OS-level bottlenecks. Argus integrates two key mechanisms: (i) a calibration methodology that involves collecting a measurement from an idle system and using it as a reference point to discover potential bottlenecks, and (ii) a tree-based data structure that represents the different OS execution paths, improving the agent's bottleneck localization accuracy. Argus aims to identify a specific kernel code path rather than stop at a subsystem-level diagnosis. In two case studies, we employ Argus to autonomously discover bottlenecks present in the memory management subsystem caused by (i) a THP aggressor co-running with other applications, and (ii) applications that incur different types of page faults. Argus produces 19 times fewer incorrect deep-path diagnoses than the strongest evaluated LLM-based baseline, which lacks reference calibration, while preserving low time-to-diagnosis (approximately 31 s)

发表机构

  • ETH Zürich(苏黎世联邦理工学院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑