ASLEval:衡量LLM智能体会话中的隐私暴露位移
ASLEval: Measuring Privacy Exposure Displacement in LLM Agent Sessions
查看机构详情
- Minjiang University(闽江学院)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
ASLEval提出隐私暴露位移概念,通过授权感知框架预先注册目标并测量所有可见出口,发现局部评估遗漏46.9%暴露,推动同时报告隐私与效用的基准。
中文摘要 AI 辅助
对使用工具的LLM智能体的隐私评估通常检查指定的动作、最终响应或攻击者报告。这些局部代理可能遗漏多步会话中其他地方的未授权暴露,并且缺乏跨出口、报告和工具路径的通用基准真相。我们引入了隐私暴露位移,即局部评估代理与目标锚定的会话暴露之间的不匹配,以及ASLEval,一个授权感知框架,该框架预先注册隐藏目标集,衡量所有声明的可见出口,并保留内部痕迹用于诊断。在多个企业风格环境和独立实现的运行时中,我们观察到三种反复出现的模式。仅期望出口的视图遗漏了可见出口并集恢复的暴露的46.9%;攻击者自我报告结合了遗漏与高错误发现率;并且模式对齐的内部证据通常在请求/探测级别先于可见暴露。减少模型可见的返回会改变这条路径,但可能消除正常任务的成功。独立的人工审查支持裁决流程,同时识别出更困难的控制台和候选案例。这些发现推动了声明完整可见边界、将主张锚定在预先指定的目标和授权中、并同时报告隐私与任务效用的基准测试。
英文摘要
Privacy evaluations of tool-using LLM agents often inspect a designated action, final response, or attacker report. These local proxies can miss unauthorized exposure elsewhere in a multi-step session and lack common ground truth across outlets, reports, and tool paths. We introduce privacy exposure displacement, the mismatch between a local evaluation proxy and target-grounded session exposure, and ASLEval, an authorization-aware framework that pre-registers a hidden target set, measures all declared visible exits, and reserves internal traces for diagnosis. Across multiple enterprise-style environments and independently implemented runtimes, we observe three recurring patterns. An expected-outlet-only view misses 46.9% of exposure recovered by the visible-exit union; attacker self-reports combine omissions with high false discovery; and schema-aligned internal evidence usually precedes visible exposure at the request/probe level. Reducing model-visible returns changes this path but can eliminate normal-task success. Independent human review supports the adjudication pipeline while identifying harder console and candidate cases. These findings motivate benchmarks that declare the complete visible boundary, ground claims in pre-specified targets and authorization, and report privacy together with task utility.