发表机构
Adobe(奥多比)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出ground-truth-as-code框架,用可执行参考函数和格式无关评判器评估实时数据科学智能体,较基线MCC提升29%,令牌消耗减少16%。
AI 中文摘要
我们提出了一个框架,用于在实时、持续更新的数据上评估数据科学智能体,该框架使用可执行的基准真值和格式无关的事实点评分。考虑这样一个示例查询:“上周的观众规模是多少”——随着底层数据的变化,参考答案也会变化,因此静态参考会过时,而标准的LLM-as-a-judge流程无法基于固定基准真值验证响应。我们的核心贡献,即“基准真值即代码”(ground-truth-as-code),将每个预期答案编码为一个可执行的参考函数,该函数在评估时直接从实时数据重新计算答案,确保参考与所描述的系统保持一致。我们将其与一个事实点级别、格式无关的评判器相结合,该评判器将智能体的响应和计算出的基准真值都分解为原子声明,并对其计算精确率、召回率和准确率,而不考虑响应格式(散文、列表、表格、HTML等)。该方法适用于预期输出可表示为可执行数据计算的智能体。我们通过一项人工与LLM一致性研究验证了该框架,该研究使用一个内部开发的、已投入生产的机器学习技能,并构建了一个合成数据库来重现生产模式和实体关系。相对于自然语言基准真值基线,我们的方法在马修斯相关系数(MCC)上实现了29%的提升——这是专家标注者与LLM-as-a-judge预测之间一致性的类平衡度量——并且每个测试用例的令牌消耗减少了16%,而缺乏明确基准真值的自导向基线与人类判断呈负相关。执行多源数据集成和非平稳数据计算的智能体在工业界被常规部署;我们提出“基准真值即代码”作为评估它们的实用方法论。
英文摘要
We present a framework for evaluating data-science agents on live, continuously updated data using executable ground truth and format-agnostic factoid scoring. Consider this example query: "what were last week's audience sizes"---the reference answer changes as the underlying data changes, so static references become outdated and standard LLM-as-a-judge pipelines cannot verify responses against a fixed ground truth. Our central contribution, ground-truth-as-code, encodes each expected answer as an executable reference function that recomputes the answer directly from live data at evaluation time, ensuring the reference remains consistent with the system it describes. We combine this with a factoid-level, format-agnostic judge that decomposes both the agent's response and the computed ground truth into atomic claims and scores precision, recall, and accuracy over them, irrespective of the response format (prose, list, table, HTML, etc.). The approach is applicable to agents whose expected outputs can be expressed as executable data computations. We validate the framework through a human--LLM agreement study on an internally developed machine learning skill deployed in production, using a synthetic database constructed to reproduce production schemas and entity relationships. Relative to a natural-language ground-truth baseline, our method achieves a 29% improvement in the Matthews Correlation Coefficient (MCC)---a class-balanced measure of agreement between expert annotators and LLM-as-a-judge predictions---and a 16% reduction in token consumption per test case, while a self-directed baseline lacking explicit ground truth is anti-correlated with human judgment. Agents that perform multi-source data integration and computation over non-stationary data are routinely deployed in industry; we propose ground-truth-as-code as a practical methodology for their evaluation.