AI 中文总结
研究如何在软件系统中可靠部署基于大语言模型的代理,提出写时复制(CoW)评分框架,利用PostgreSQL级机制在应用环境中直接评估代理操作,能低成本评估迭代,在开源平台上验证了该框架的有效性。
AI 中文摘要
在软件系统中可靠地部署基于大语言模型的代理需要评估它们在特定应用工作流程中的表现,且要有足够的粒度来定位其成败之处。然而,现有的代理评估机制存在局限性:基准测试对特定应用工作流程和环境的结构效度较低,复制评估环境成本高昂且容易出现偏差。我们提出了写时复制(CoW)评分框架,该框架使用PostgreSQL级别的写时复制机制在应用环境中直接评估代理操作,以隔离代理写入。CoW评分产生会话级和操作级分数,突出代理数据库写入操作在给定应用环境中的成败之处,从而能够对代理工具和工具表面进行低成本评估和迭代。我们在开源项目管理平台Plane上展示了该框架,分析揭示了工具表面的特定问题,相应的修复措施对受影响的模型产生了可衡量的改进。
英文摘要
Trustworthy deployment of LLM-based agents in software systems requires evaluating how they perform on application-specific workflows, with enough granularity to localize where they succeed and fail. Yet existing agent evaluation mechanisms are limited: benchmarks have low construct validity for application-specific workflows and environments, and replica evaluation environments are expensive and prone to drift. We propose Copy-on-Write (CoW) Scoring, a framework that evaluates agent operations directly within application environments using a PostgreSQL-level Copy-on-Write mechanism to isolate agent writes. CoW Scoring produces session- and operation-level scores that highlight where agents' database write operations succeed and fail in a given application environment, enabling inexpensive evaluation and iteration on agent harnesses and tool surfaces. We demonstrate the framework on Plane, an open-source project-management platform, where analysis surfaced specific issues in the tool surface, and corresponding fixes produced measurable improvements on affected models. Python library: https://github.com/trail-ml/agent-cow-python
Comments15 pages, 11 figures, accepted at ICML 2026 Second Workshop on Agents in the Wild: Safety, Security, and Beyond