Transect:为长时程LLM智能体评估保留可观测性
Transect: Retaining Observability for Long-Horizon LLM Agent Evaluations
浏览论文内容
中文总结 AI 辅助
Transect是一个基于Inspect Scout的开源工具,通过可导航报告和可追溯标签,帮助评估者理解长时程智能体运行,保留可观测性,支持科学严谨性和可复现性。
中文摘要 AI 辅助
前沿AI评估越来越多地使用开放式、智能体化、长时程任务,其转录内容可能涵盖复杂多智能体网络中数百页的输出和行动。因此,可观测性范围——评估者能够可靠推断智能体行为的范围——正在缩小。语言模型助手可以帮助分类和解释智能体行为,但也赋予人类评估者显著的分析自由度,威胁到基于语言模型的转录分析的可复现性和可审计性。Transect是一个基于Inspect Scout构建的开源软件包,旨在帮助评估者理解长智能体运行如何展开,识别值得调查的行为,并对照转录检查解释。用户可在可复用的评估家族配置中指定任务上下文和行为词汇表,而评判模型和分析设置则单独提供。Transect的可导航报告将记录的事件、令牌使用、子智能体活动和模型生成的行为标签对齐在共同的基于回合的时间线上。审查者可以快速掌握运行的叙事,将任何标签或事件追溯到其来源回合,并导出底层数据表以进行跨运行分析。我们在一个AI研发评估中演示了该工作流程,该评估生成了近1300万令牌,将智能体的工作划分为与研究技能分类、子智能体委派和交互以及令牌使用对齐的行为阶段。综合视图显示了对操作工作和手稿制作的关注,而几乎没有持续假设生成阶段的证据——这可以说是高质量科学输出的必要组成部分。Transect灵活、可定制的转录分析管道将使评估者能够跟上更长、更复杂、更频繁的AI评估,同时支持科学严谨性、透明度和可复现性。
英文摘要
Frontier AI evaluations increasingly use open-ended, agentic, long-horizon tasks whose transcripts can span hundreds of pages of outputs and actions from complex multi-agent networks. The observability envelop-the range of what evaluators can reliably infer about an agent's behaviours-is therefore narrowing. Language model assistants can help classify and interpret agent behaviour but also afford human evaluators significant analytical degrees of freedom, threatening the reproducibility and auditability of language-model-based transcript analysis. Transect is an open source package built on Inspect Scout to help evaluators understand how a long agent run unfolded, identify behaviour worth investigating, and check interpretations against the transcript. Users specify task context and behavioural vocabulary in a reusable evaluation-family configuration, with judge models and analysis settings supplied separately. Transect's navigable reports align recorded events, token use, sub-agent activity, and model-generated behavioural labels on a common turn-based timeline. Reviewers can quickly grasp a run's narrative, trace any label or event to its source turns, and export the underlying data tables for cross-run analysis. We demonstrate the workflow on an AI R&D evaluation that generated almost 13 million tokens, dividing the agents' work into behavioural phases aligned with research-skill classifications, sub-agent delegations and interactions, and token use. The combined view shows a focus on operational work and manuscript production, with little evidence of a sustained hypothesis generation stage-arguably a necessary component for high-quality scientific outputs. Transect's flexible, customisable transcript-analysis pipeline will enable evaluators to keep pace with longer, more complex, more frequent AI evaluations while supporting scientific rigour, transparency, and reproducibility.
发表机构
- UK AI Security Institute(英国人工智能安全研究所)
- Meridian Labs(Meridian 实验室)
机构由 AI 辅助整理,请以论文原文为准。