arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

标签并非终点:MCP智能体安全评估中的处理泄露与构念效度

Labels Are Not Endpoints: Treatment Leakage and Construct Validity in MCP Agent Security Evaluation

Rana Muhammad Ahmed, Sabahat Abbas

arXiv 2608.12880首次发表:更新:

发表机构

Bahria University(巴里亚大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究针对MCP智能体安全评估中标签与行为事实的偏差问题,通过审计留存评估活动发现处理泄露现象,提出七环节完整性链及端点完整性检查器,完成活动受限的测量审计。

AI 中文摘要

对使用工具的智能体进行安全评估时,常将存储的标签等同于行为事实。我们通过追溯10200条执行行至180个模型绑定请求、45个语义请求及15个可观测刺激,对一项已留存的评估活动展开审计。本次审计实施了两类处理方案,但未用到计划中的外部有效载荷系列语料库。历史 grader 表现出直接的处理泄露:处理元数据对ATTACK_SUCCESS类进行了门控,因此固定行为在处理重标注下可改变类别。一项处理盲法重构将58条历史ATTACK_SUCCESS或HIJACK_ATTEMPT标签修正为授权的良性完成内容,同时保留了3个经核实的受保护数据传输及1个单独的未授权转发案例。锁定的v2普查中恰好无ATTACK_SUCCESS记录,而该转发案例仍为关于目标完成的语义边界处的HIJACK_ATTEMPT。对锁定v2判定为可结构解释的全部96个请求开展双评审员盲法一致性审查,产生了评审员共识类别,但在4个构念边界案例上与锁定代码手册存在差异。我们提出了七环节完整性链及一个可执行、范围受限的端点完整性检查器。最终成果是一项活动受限的测量审计,而非总体攻击率、模型排名、防御效能或因果估计。

英文摘要

Security evaluations of tool-using agents often equate stored labels with behavioral facts. We audit a preserved campaign by tracing 10,200 execution rows to 180 model-bound requests, 45 semantic requests, and 15 observable stimuli. Two schema treatments were delivered, but the planned external payload-family corpus was not. The historical grader exhibited direct treatment leakage: treatment metadata gated the ATTACK_SUCCESS class, so fixed behavior could change class under treatment relabeling. A treatment-blind reconstruction corrects 58 historical ATTACK_SUCCESS or HIJACK_ATTEMPT labels to authorized benign completions while preserving three verified protected-data transfers and one separate unauthorized-forwarding case. The locked v2 census contains exactly zero ATTACK_SUCCESS records, while the forwarding case remains a HIJACK_ATTEMPT at a semantic boundary concerning objective completion. A dual-reviewer blinded concordance review of all 96 requests deemed structurally interpretable by locked v2 produced identical reviewer-consensus classes but differed from the locked codebook on four construct-boundary cases. We contribute a seven-link Integrity Chain and an executable, scope-bounded endpoint-integrity linter. The result is a campaign-bounded measurement audit, not a population attack-rate, model-ranking, defense-efficacy, or causal estimate.

Comments24 pages, 10 figures, 4 tables. Preprint

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑