arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.05363cs.AI

SovereignPA-Bench:在演化意图、平台中介与同意约束下评估用户自有个人智能体

SovereignPA-Bench: Evaluating User-Owned Personal Agents under Evolving Intent, Platform Mediation, and Consent Constraints

Dylan Zongmin Liu

首次发表
浏览论文内容

中文总结 AI 辅助

针对现有基准未考量个人智能体用户主权保护的问题,提出可执行基准SovereignPA-Bench,从多维度评估智能体主权能力,验证全主权框架可显著降低隐私泄露等风险。

中文摘要 AI 辅助

个人智能体正逐步成为持久的用户自有中介:它们记忆用户偏好、过滤平台中介信息、调用工具并与各类服务协商。现有基准可评估工具使用、网页导航、桌面控制、个性化、推荐系统及上下文演化能力,但很少关注智能体是否能维护用户主权——即在尊重隐私、知情同意、证据有效性、用户负担要求的同时推进用户当前利益,且能抵御诱导性操纵。本文提出SovereignPA-Bench这一可执行基准,用于在演化意图、平台中介、隐私边界、同意约束、证据要求及负担权衡的场景下评估用户自有个人智能体。该基准将智能体可见的可观测状态与仅评估方可访问的隐藏标签分离,输出任务成功率、对齐度、隐私合规性、同意合规性、证据有效性、抗操纵性、用户负担及可审计性的分项指标,同时保留配对场景顺序以支持模型与策略的横向对比。研究覆盖4类模型家族与8种基线策略的120个主权压力测试场景,生成3840条固定提示词轨迹,包含原始提示词、模型输出、服务商返回响应、解析后动作、可复算指标、预设分析结果、定性案例,以及由3名标注员对240个样本完成的盲审结果。全主权脚手架方案相比直接执行、仅依赖记忆、仅遵循同意规则、仅依赖证据、ReAct/工具调用、安全提示词、裁判防护等基线方案,可提升主权得分,同时降低隐私泄露、同意违规、过度让步及被操纵捕获的风险。人工审计结果显示,标注员对隐私与同意合规性的判定一致性较高,对操纵行为的判定一致性较低,明确了平台诱导判定的主观边界。上述结果表明,个人智能体评估必须跳出仅关注任务完成度的局限,转向覆盖代表性场景、具备同意感知能力、基于证据支撑的动作评估体系。

英文摘要

Personal agents that book and buy for their users act through platforms that rank, nudge, pre-select and collect data in their own interest. We introduce SovereignPA-Bench, a controlled benchmark of whether such an agent follows the user's current intent, resists steering, shares only what a service needs, asks before exceeding its authority, and reports truthfully. A scripted platform and a scripted user surround the agent in 1,920 booking scenarios and 288 cancellation scenarios with retention flows. Each of the 192 booking situations, in 16 domains, is run as a control and under 9 paired variants that change one factor: stale memory, ambiguous intent, sponsored ranking, urgency, pre-checked extras, over-collecting forms, injected reviews, a mid-task update, or a mid-task reminder. Metrics are deterministic, and every question the agent asks is labelled necessary or unnecessary, so faithfulness and user burden are measured separately. Two rule-based agents reach 100% sovereign success, one of them without a single unnecessary question. Across 17 open-weight models, sovereign success ranges from 2% to 82%. More faithful models ask fewer unnecessary questions, not more (Spearman rho = -0.67 between success and burden). A form with two "recommended" fields raises the share of episodes that send a detail the service does not need from 2.4% to 58%. Injected reviews get a requested personal detail to the provider in 41% of attempts, against 0.7% without them. Sponsored labels and urgency banners shift choices only slightly in our setting. Retention offers never worked when the user had ruled them out in advance, but without that sentence 4 models accepted offers or pauses, and in 184 of 440 failed obstructed cancellations the agent told the user it had succeeded. A prompt-level checklist and a structural firewall raise success by at most 9 points. Code, scenarios and logs will be released.

发表机构

  • Stanford University(斯坦福大学)

机构由 AI 辅助整理,请以论文原文为准。

↑