AgentWorld:面向智能体信息检索的人格感知可靠性评估
AgentWorld: Personality-Aware Reliability Evaluation for Agentic Information Retrieval
浏览论文内容
中文总结 AI 辅助
AgentWorld框架整合多模块,通过三类实验验证其可评估智能体信息检索的人格感知可靠性,能发现统一测试未暴露的故障模式及轨迹级脆弱性。
中文摘要 AI 辅助
智能体信息检索的评估目前仅限于与统一用户的脚本化交互,缺失了自然人格多样性与对抗性脆弱性。我们提出AgentWorld,这一模拟框架整合了四部分内容:(i)基于大五人格(OCEAN)的用户群体与有状态工具使用环境;(ii)具备结构化故障分类、部分计分及双控制交接验证的pass^k一致性度量;(iii)六种微调格式下的分数阈值训练数据导出;(iv)对抗性风险分析器,该分析器会快照所需的中间状态主干,在四种任务感知扰动类型下分支蒙特卡洛滚动,并通过ΔP/ΔT计分、Dempster–Shafer证据融合及Shapley攻击类别归因来量化风险。三项实验验证了该框架:针对10种OCEAN人格的对话分析智能体(240条评估者判断);针对5项任务×4种人格变体的客服智能体;以及对5项任务的对抗性压力测试,揭示了存在的轨迹脆弱性(无扰动时V_min=0.375)及工具/基础设施层攻击的主导性(Shapley值:系统层46%,动作层38%)。人格差异会暴露统一测试无法发现的故障模式——跨域泄漏、上下文漂移、0.27分的质量差距,以及同一任务下不同人格的通过率分别为50%与100%;而风险分析器能量化仅靠pass^k无法测量的轨迹级脆弱性。
英文摘要
Evaluation of agentic information retrieval remains limited to scripted interactions with uniform users, missing both natural personality diversity and adversarial brittleness. We present AgentWorld, a simulation framework combining (i)Big Five (OCEAN) personality-driven user populations with stateful tool-use environments; (ii)the pass$^k$ consistency metric with structured fault classification, partial-credit scoring, and dual-control handoff verification; (iii)score-thresholded training-data export in six fine-tuning formats; and (iv)an adversarial Risk Analyser that snapshots required-intermediate-state spines, branches Monte-Carlo rollouts under four task-aware perturbation types, and quantifies risk via $ΔP / ΔT$ scoring, Dempster--Shafer evidence fusion, and Shapley attack-category attribution. Three experiments demonstrate the framework: a conversational analytics agent across 10 OCEAN personas (240 evaluator judgments); a customer-support agent across 5 tasks $\times$ 4 persona variants; and adversarial stress-testing of 5 tasks revealing pre-existing trajectory brittleness ($V_{\min}=0.375$ without perturbation) and tool/infrastructure-layer attack dominance (Shapley: 46% system, 38% action). Personality variation surfaces failure modes uniform testing cannot expose---cross-domain leakage, contextual drift, a 0.27-point quality gap, and 50% vs. 100% pass-rate across personas on the same task---while the Risk Analyser quantifies trajectory-level brittleness that pass$^k$ alone cannot measure.
发表机构
- PayPal(贝宝)
机构由 AI 辅助整理,请以论文原文为准。