arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

BTS-AgentBench:从只读遥测日志到智能体基准的确定性可重放流水线

BTS-AgentBench: A Deterministic, Replayable Pipeline from Read-Only Telemetry Logs to Agent Benchmarks

Jeong-Yoon Kim

arXiv 2608.27334首次发表:更新:

AI 中文总结

BTS-AgentBench 提供从只读遥测日志到可执行多轮智能体任务的确定性可重放流水线,经多轮验证可复现基准,应用于 XAI4HEAT 生成高质量任务,为智能体基准构建提供标准化方案。

AI 中文摘要

工业站点包含大量只读遥测数据,但很少有基准明确说明如何将这些记录编译为可执行的多轮智能体任务。我们提出一种遥测到 episode(任务片段)的构建方法,实例化为 BTS-AgentBench。该流水线将 BTS 元数据和原始历史归一化为只读工具库,编译带有工具衍生的标准答案和证据的静态任务,并将保留的任务提升为类型化、有界的面向操作符的 episode。此次 532 行的发布补充了澄清、目标修订、时间戳策略、质量门控报告和证据归因,同时保留了源计算和划分。编码的合同预检查报告零发现,构建排除控制器完成 0/532 行。两次独立的原始到 episode 构建匹配全部 11 个逻辑工具库导出,并精确复现发布的 356/87/89 训练/验证/测试工件。将该共享构建路径应用于 XAI4HEAT 生成 204 个 episode;在其 41 行保留的测试划分上,控制器完成 0 行,保留的 GPT-5.5 执行完成全部 41 行。代码、工件和重放报告可在此 https URL 获取。

英文摘要

Industrial sites contain large volumes of read-only telemetry, but few benchmarks specify how to compile these records into executable multi-turn agent tasks. We present a telemetry-to-episode construction method instantiated as BTS-AgentBench. The pipeline normalizes BTS metadata and raw histories into a read-only tool store, compiles static tasks with tool-derived gold answers and evidence, and lifts retained tasks into typed, bounded operator-facing episodes. The 532-row release adds clarification, goal revision, timestamp policy, quality-gated reporting, and evidence attribution while preserving the source computation and split. Coded contract preflight reports zero findings, and the construction-exclusion controller completes 0/532 rows. Two independent raw-to-episode builds match all 11 logical tool-store exports and reproduce the released 356/87/89 train/dev/test artifact exactly. Applying the shared construction path to XAI4HEAT produces 204 episodes; on its 41-row held-out test split, the controller completes 0 rows and the retained GPT-5.5 execution completes all 41. Code, artifacts, and replay reports are available at https://github.com/kjy7567/BTS-AgentBench.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑