AI 中文总结
该研究提出Spike-Killer,一种人在环的证据门控工作流,用于安全诊断Windows工作站的帧时间问题,以CS2为目标应用验证了其可审计性,为LLM辅助智能体提供了可信操作模式。
AI 中文摘要
基于大语言模型(LLM)的智能体可整合系统证据、提出配置变更方案并自动化诊断任务,但其灵活性会导致操作风险,例如动作不精确或收集程序具有侵入性。本文提出Spike-Killer,一种经人工批准的工作流,用于在真实Windows工作站上诊断帧时间投诉问题。该工作流将每个动作视为证据门控事务:记录精确的目标状态、对风险进行分类、保留快照、验证后置条件,并将失败的测量值作为一级证据保留。本文为一篇经验论文,报告了一项当日完成的研究,以《反恐精英2》(Counter-Strike 2)作为高要求目标应用程序。证据包包含保留的状态快照、探索性微基准测试、十次相同状态下的可重复性探测、实时遥测、修复的过宽注册表操作、不兼容的画面捕获尝试、无效的本地回放以及系统级跟踪替换。Windows性能记录器(Windows Performance Recorder)生成了两条CS2本地机器人GPU跟踪,时长分别为90.69秒和85.85秒;两者均归因于该http URL,暴露了DxgKrnl Present元数据,且零丢失ETW缓冲区或事件。这些结果验证了跟踪完整性,而非性能:本研究未报告任何帧间隔、P99估计值或干预效果。本研究的贡献是一种可审计的人在环模式,用于在真实工作站上实现可信的智能体辅助,包括证据不足时的明确停止条件。
英文摘要
LLM-assisted agents can synthesize system evidence, propose configuration changes, and automate diagnostic tasks, but their flexibility makes an imprecise action or an intrusive collector an operational risk. We present Spike-Killer, a human-approved workflow for diagnosing frame-time complaints on one real Windows workstation. The workflow treats each action as an evidence-gated transaction: it records the exact target state, classifies risk, preserves a snapshot, verifies a postcondition, and retains failed measurements as first-class evidence. This experience paper reports a completed same-day study with Counter-Strike 2 as a demanding target application. The evidence bundle contains preserved state snapshots, exploratory microbenchmarks, a ten-run same-state repeatability probe, live telemetry, a repaired over-broad registry action, incompatible presentation-capture attempts, an invalid local replay, and a system-level tracing replacement. Windows Performance Recorder produced two CS2 local-Bot GPU traces of 90.69 and 85.85 seconds; both were attributed to cs2.exe, exposed DxgKrnl Present metadata, and had zero lost ETW buffers or events. These results qualify trace integrity, not performance: the study reports no frame intervals, P99 estimate, or intervention effect. The contribution is an auditable, human-in-the-loop pattern for trustworthy agent assistance on a real workstation, including explicit stop conditions when evidence is insufficient.
Comments6 pages. Experience report; no causal CS2 frame-time improvement or autonomous-safety claim is made