arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.02975cs.AIcs.CL

使用Incognita评估基于社会分布式任务环境中行动的生成智能体

Where Knowledge and Authority Sit Changes What an Agent Benchmark Can Resolve

  • RedMind Research(RedMind研究)

机构由 AI 辅助整理,请以论文原文为准。

Dan C. Hsu, Luke Lu

AI总结:

研究结合现有基准和社会模拟环境要求的评估设置,定义社会分布式任务环境,介绍Incognita框架,以之评估三个生成智能体模型,发现其在奖励和行为上有进展但可靠性仍低。

AI中文摘要:

社会环境中的有效智能体取决于何时寻求知识、行动以及行动是否合理。现有基准提供可执行行动等,社会模拟环境提供丰富交互。我们研究结合这些要求的评估设置,定义社会分布式任务环境,引入Incognita框架,评估三个生成智能体模型,发现有进展但可靠性低。

英文摘要:

Most agent benchmarks put facts, tools and permissions behind one interface. Real organizations spread them across people. Incognita asks what happens when the task and success criterion stay fixed but access does not. We transform eighteen customer-service tasks into three settings: direct access, one known intermediary, and six role-isolated participants whose capabilities must be discovered. Across 864 trials with four models, social access reduced success for every model; the pre-specified intervals excluded zero for two. The latest tested model, gpt-5.6-sol, achieved the highest social-access success at 0.65, a 0.11 decrease from centralized indirect access with an interval that included zero. Exploratory comparisons separated five of six model pairs under social access, while neither centralized setting separated any at this sample size. In post-hoc task-blocked tests, three pairwise interaction $p$-values remained significant after multiplicity adjustment. A reference-relative reader associates the wider gaps with failures to obtain needed information. Because the data cannot distinguish ineffective requests by the evaluated agent from inaccurate replies by simulated participants, the reader's labels describe where trajectories stopped, not why models differed.

↑