AI 中文总结
本研究通过两个实验揭示LLM金融智能体评估中执行比较与审计分数的测量边界,表明固定磁带执行和诊断控制(如多缺陷审计)对分数解释至关重要。
AI 中文摘要
在LLM智能体评估中,需要哪些控制措施来解释执行性能和审计分数?我们在一个金融智能体框架中研究了这些解释的两个限制。在研究A中,比较了在理想化和压力执行条件下三个共享一个24天上涨阶段的合成设置上的独立运行,将执行规则与新的模型响应和投资组合反馈混合在一起:解析后的决策路径在450对中仅有19.8%一致。将每个存储的响应磁带重放到两个执行目的地会得到更窄的结果。以这些响应为条件,压力执行将总回报改变-0.0170(95%区间[-0.0230,-0.0117]),即理想化基线的10.4%,且十个种子聚类无法解决模型排名问题。研究B纠正了一个不完整的答案键,并将遗留任务替换为在显式多标签提示下的匹配的零缺陷、单缺陷和双缺陷任务。从单缺陷到双缺陷,目标违规召回率的下降在审计器和来源的六种组合中有五种为正值(中位数0.267),其中三种在Holm校正后仍然显著。然而,包含两个目标标签的审计器最常出现微精度0.149,在98/100个零缺陷任务上发出发现,并且仅在21/100个案例中返回精确的双缺陷集。因此,单独的目标召回率在这种构造上无法很好地说明审计质量。这些研究解决了不同的限制:执行比较估计了什么,以及目标召回率捕获了什么。它们共同表明,固定条件和诊断控制如何限制一个分数所能支持的断言。
英文摘要
What controls are needed to interpret execution performance and audit scores in LLM agent evaluations? We study two limits on these interpretations in a financial agent harness. In Study~A, comparing independent runs under idealized and stressed execution on three synthetic settings that share one 24-day upward phase mixes the execution rule with fresh model responses and portfolio feedback: the parsed decision paths agree in only $19.8\%$ of $450$ pairs. Replaying each stored response tape through both execution destinations gives a narrower result. Conditional on those responses, stressed execution changes total return by $-0.0170$ (95\% interval $[-0.0230,-0.0117]$), or $10.4\%$ of the idealized baseline, and ten seed clusters do not resolve the model ranking. Study~B corrects an incomplete answer key and replaces legacy tasks with matched zero-, one-, and two-defect tasks under an explicit multi-label prompt. The drop in target violation recall from one to two defects is positive in five of six combinations of auditor and source (median $0.267$), with three surviving Holm correction. Yet the auditor that includes both target labels most often has micro-precision $0.149$, emits findings on $98/100$ zero-defect tasks, and returns the exact dual-defect set in only $21/100$ cases. Target recall by itself therefore gives a poor account of audit quality on this construction. The studies address different limits: what an execution comparison estimates, and what target recall captures. Together, they show how fixed conditions and diagnostic controls bound the claims a score can support.