AI 中文总结
针对生成式大模型评判智能体安全存在延迟高、输出易失效的问题,该研究提出用类型化决策模型JEV做回溯轨迹分类,实验证明其在多基准上整体F1更高、成本更低,可作为经济型安全筛选方案。
AI 中文摘要
使用工具的智能体的安全评估需要结合上下文判断行为,但生成式评判器会带来延迟、解释开销以及输出验证失败的问题。我们研究了JEV(一种类型化决策模型)是否能为回溯式轨迹分类提供有效的替代方案。我们在共包含5219条轨迹的四个基准集合上,使用统一的风险评分标准和行为级标签,评估了JEV与四款生成式评判器的表现。JEV的基准平均正类F1值达到77.8,而表现最强的生成式配置GLM-5.2为74.1;二者的有效结果覆盖率分别为95.5%和94.4%。不同数据集上的表现存在差异:JEV在ATBench500和MCPHunt上领先,GLM在R-Judge和TraceSafe上领先。在四个基准测试中,JEV的成功调用中位延迟为0.99秒;每次有效判断的预估token成本平均为0.000195美元。这些结果表明JEV可作为一种经济的筛选信号,但在精确率和召回率上存在权衡。
英文摘要
Security evaluation of tool-using agents requires judging actions in context, yet generative judges add latency, explanation overhead, and output-validation failures. We study whether JEV, a typed decision model, offers a useful alternative for retrospective trace classification. We evaluate JEV and four generative judges on four benchmark collections totaling 5,219 trajectories, using a common risk rubric and behavior-level labels. JEV attains a benchmark-averaged positive-class F1 of 77.8, compared with 74.1 for the strongest generative configuration, GLM-5.2, with valid-result coverage of 95.5\% and 94.4\%, respectively. Performance varies across datasets, with JEV leading on ATBench500 and MCPHunt and GLM leading on R-Judge and TraceSafe. Across the four benchmarks, JEV's median successful-call latency is 0.99 seconds; estimated token cost averages \$0.000195 per valid judgment. These results support JEV as an economical screening signal, with trade-offs in precision and recall.