GuardianAgent:基于策略的风险自适应匿名化与经验证的对抗升级
GuardianAgent: Policy-Conditioned Risk-Adaptive Anonymization with Verified Adversarial Escalation
浏览论文内容
中文总结 AI 辅助
GuardianAgent是一种基于策略的匿名化框架,通过AMRSF计算风险、采用快慢路径机制与五级重写层级,在三类基准上实现最优隐私-效用权衡,且对不同上下文条件决策具有鲁棒性。
中文摘要 AI 辅助
实时网络流量的隐私保护仅检测隐私片段是不够的,基于智能体的隐私保护系统必须判断输出操作是否符合目标网站的隐私政策,然后仅应用与剩余泄露风险相匹配的重写或净化级别。我们提出了GuardianAgent,这是一种基于策略的匿名化框架,将结构化风险评估与经验证的自适应重写相结合。GuardianAgent通过AMRSF(自适应多因素风险评分公式)计算风险,这是一个明确的控制器,将政策违规可能性与数据敏感性、接收方传输、目的合法性、上下文依据和政策透明度相结合,而非依赖LLM直接分配风险。该风险评分决定允许/转换/拒绝的决策以及初始匿名化级别。为提高效率,GuardianAgent对低不确定性的策略匹配使用证据快速路径,仅对不确定情况调用LLM慢速路径。对于重写,它应用由经验证的对抗猜测器驱动的五级层级:仅当猜测得到原始文本支持时才触发升级,防止幻觉的攻击者信心导致不必要的过度匿名化。在涵盖法律文本(TAB)、Reddit帖子(SynthPAI)和多格式合成PII记录(PII-Masking-300k)的三个基准上的实验表明,GuardianAgent在已发布基线中实现了最强的隐私-效用权衡,是唯一在所有三个领域均达到超过0.90隐私的方法,且在骨干切换下保持鲁棒性。操作上下文压力测试进一步显示,相同的输出文本在不同接收方、目的、操作依据和政策透明度条件下会收到不同的决策和匿名化强度。
英文摘要
Privacy protection for live web traffic requires more than detecting private spans. Agent-based privacy protection systems must determine whether an outgoing action complies with the destination site's privacy policy, then apply only the level of rewriting or sanitisation justified by the residual disclosure risk. We present GuardianAgent, a policy-conditioned anonymization framework that couples structured risk assessment with verified adaptive rewriting. GuardianAgent computes risk through AMRSF (Adaptive Multi-factor Risk Scoring Formula), an explicit controller that combines policy-violation likelihood with data sensitivity, recipient transmission, purpose legitimacy, contextual basis, and policy transparency, rather than relying on an LLM to assign risk directly. This risk score determines both the allow/transform/deny decision and the initial anonymization level. For efficiency, GuardianAgent uses an evidential fast path for low-uncertainty policy matches and invokes an LLM slow path only for uncertain cases. For rewriting, it applies a five-level hierarchy driven by a verified adversarial guesser: guesses trigger escalation only when supported by the original text, preventing hallucinated attacker confidence from causing unnecessary over-anonymization. Experiments across three benchmarks spanning legal text (TAB), Reddit posts (SynthPAI), and multi-format synthetic PII records (PII-Masking-300k) show that GuardianAgent achieves the strongest privacy-utility trade-off among published baselines and is the only method to reach more than 0.90 privacy in all three domains, remaining robust under a backbone switch. Action-context stress tests further show that the same outgoing text receives different decisions and anonymization strengths under different recipients, purposes, action bases, and policy-transparency conditions.
发表机构
- School of Computer Science and Engineering, UNSW Sydney(新南威尔士大学悉尼分校计算机科学与工程学院)
机构由 AI 辅助整理,请以论文原文为准。