arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MemRiskBench:面向长时程LLM智能体的轨迹感知风险保持评估

MemRiskBench: Trace-Aware Risk-Preserving Evaluation for Long-Horizon LLM Agents

Jianhua Jiang, Dongbo Yuan, Weihua Li

arXiv 2609.14976首次发表:更新:

发表机构

School of Artificial Intelligence and Computer Science, Jilin University of Finance and Economics; Jilin Province Key Laboratory of Fintech, Jilin University of Finance and Economics; School of Engineering, Computer and Mathematical Sciences, Auckland University of Technology(吉林财经大学人工智能与计算机科学学院; 吉林省金融科技重点实验室,吉林财经大学; 奥克兰理工大学工程、计算机与数学科学学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

MemRiskBench提出五类风险分类与轨迹接地检查的120片段基准,并设计风险保持子集选择器,在20%子集下保持排序与高风险检测,计算量减少5倍。

AI 中文摘要

长时程LLM智能体会在会话间累积记忆,从而产生稀疏但高影响的风险:过时事实、冲突更新、跨用户泄露、撤销记忆的复用以及约束衰减。标准的聚合分数掩盖了各类风险的失败率——一个平均准确率达到78%的模型仍可能在4%的片段中泄露数据——而基准压缩则优先丢弃那些罕见的、高严重性的事件,这些事件恰恰区分了基本可用的模型与偶尔造成危害的模型。我们提出了MemRiskBench。主要贡献是:第一,一个五类别风险分类体系(外加一个已记录但不计分的类别),通过确定性轨迹接地检查实现,实例化为一个包含120个片段的脚本化基准,具有完整的轨迹日志记录,且在通过/失败判定路径上不使用LLM作为裁判,并在五个本地运行的量化指令微调模型上进行了评估。第二,一个风险保持子集选择器:基于确定性轨迹衍生特征的覆盖率约束贪心选择器,在20%的子集规模下保持完整排序(Spearman rho = 0.975,确定性;置信区间因零自举方差而坍缩为点估计)、风险覆盖率(1.0)和高风险模型检测率(1.0),并将计算量减少5倍。与仅关注排序的子集选择器不同,该选择器还通过基于轨迹的确定性特征(无需LLM裁判)保持风险类型覆盖和高风险检测。所有片段、轨迹、评分实现和选择器均已发布,以支持对已部署LLM智能体的可复现评估和风险评估。

英文摘要

Long-horizon LLM agents accumulate memory across sessions, creating sparse but high-impact risks: stale facts, conflicting updates, cross-user leakage, revoked-memory reuse, and constraint decay. Standard aggregate scores hide per-risk failure rates--a model achieving 78% average accuracy may still leak data in 4% of episodes--and benchmark compression preferentially discards the rare high-severity events that distinguish a mostly-working model from one that occasionally causes harm. We present MemRiskBench. The primary contribution is a five-category risk taxonomy (plus one documented, unscored category) operationalized by deterministic trace grounded checks, instantiated as a 120-episode scripted benchmark with full trace logging and no LLM-as-judge on the pass/fail path, evaluated on five locally run quantized instruction-tuned models. Second, a risk-preserving subset selector: a coverage-constrained greedy selector on deterministic trace-derived features that retains full ranking (Spearman rho = 0.975, deterministic; CI collapses to a point estimate with zero bootstrap variance), risk coverage (1.0), and high-risk model detection (1.0) at a 20% subset size, reducing compute 5x. Unlike ranking-only subset selectors, this selector additionally preserves risk-type coverage and high-risk detection using trace-grounded deterministic features that do not require an LLM judge. All episodes, traces, the scoring implementation, and the selector are released to support reproducible evaluation and risk assessment of deployed LLM agents

Comments11 pages, 4 figures,

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑