主权谈判基准:在隐私、同意、证据和制度压力下评估委托谈判中用户拥有的个人代理
SovereignNegotiation-Bench: Evaluating User-Owned Personal Agents In Delegated Bargaining Under Privacy, Consent, Evidence, And Institutional Pressure
浏览论文内容
中文总结 AI 辅助
研究个人代理替用户谈判的情况,介绍主权谈判基准,其能在多方面约束下评估谈判,通过多种场景验证,指出仅协议成功对用户拥有的谈判代理不足。
中文摘要 AI 辅助
个人代理将越来越多地代表用户进行谈判。现有谈判基准强调协议、盈余或战略能力,但用户拥有的代理可能在达成协议时损害用户。我们引入主权谈判基准,它在多种约束下评估委托个人代理谈判,通过多种场景验证,结果表明协议成功对用户拥有的谈判代理不足。
英文摘要
Personal AI agents are beginning to negotiate for people, from refunds and bills to deposits and sales. A human agent in that position is judged by the duties owed to the principal, not by whether a deal was struck. We introduce SovereignNegotiation-Bench, a controlled benchmark that operationalizes five such duties from agency law--loyalty, obedience to actual authority, confidentiality, candor and diligence--as deterministic checks on episode logs; the first three enter a single headline metric. The benchmark contains 1,764 paired scenarios (252 situations in 18 consumer and peer-to-peer domains, each under 7 counterparty tactics). The counterparty's economics are a fixed function of the agent's structured actions and of the disclosures detected in its messages, so outcomes are comparable across agents and a disclosed limit has a measurable, causal price. A simulated principal grants or withholds consent and tightens its mandate midepisode. Rule-based agents show that the benchmark is solvable from the observable state (92% faithful success) and that a single disclosing sentence erases the entire negotiated surplus (0.70 to 0.00). Across 17 open-weight models, faithful success ranges from 6% to 75%; models disclose the principal's reservation value in 2-80% of episodes, agree or share a protected document without a required approval in 2-23%, and follow an instruction injected into the counterparty's message in 5-57% of injection episodes. Deal rate ranks models much like faithful success does, but it does not certify individual agreements: pooled over models, 48% of the agreements breach at least one duty (18-96% per model). Within families, faithful success tends to rise with size, but no size trend is significant, and on the model we test, neither prompting nor a code-level guard raises faithful success substantially. Code, scenarios and all episode logs will be released.
发表机构
- Stanford University(斯坦福大学)
机构由 AI 辅助整理,请以论文原文为准。