发表机构
Stanford University; University of British Columbia(斯坦福大学; 不列颠哥伦比亚大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出一种结合规则市场动态与LLM驱动代理的模拟架构,以分析AI评估生态系统,并发现基准保留设计对评估分数与用户满意度差距的影响具有异质性。
AI 中文摘要
AI评估影响着模型提供者、用户、资助者和监管者的决策。我们认为,设计有效的基准需要将设计选择置于该参与者生态系统的动态中进行情境化。我们开发了一种模拟架构,该架构将基于规则的市场动态与LLM驱动的战略参与者相结合,基于生成式基于代理建模(GABM)的进展。我们将基准、消费者需求和提供者能力建模为六维能力空间(推理、编码、知识、安全、沟通、智能体)上的向量,并在参与者之间设置结构性信息分区。作为案例研究,我们将这种风格化模拟应用于探索基准保留(holdout)设计。我们发现,从公共基准转向私有保留基准,在大多数基准上缩小了基准分数与用户满意度之间的差距,但在少数基准上扩大了差距,这取决于保留权重将评分信用转移至何处。我们分别在工具和案例研究层面对我们的发现进行了压力测试,借鉴了Sargent(2013)的V&V框架和GABM特定的证据标准。除了保留设计之外,我们的模拟是一个假设生成的沙盒,用于研究评估者和政策选择如何反过来重塑生态系统。
英文摘要
AI evaluation shapes the decisions of model providers, users, funders, and regulators. We argue that designing valid benchmarks requires contextualizing design choices in the dynamics of this ecosystem of actors. We develop a simulation architecture that combines rule-based market dynamics with LLM-driven strategic actors, building on advances in Generative Agent-Based Modeling (GABM). We model benchmarks, consumer needs, and provider capabilities as vectors over a six-dimensional capability space (reasoning, coding, knowledge, safety, communication, agentic), with structural information partitions across actors. As a case study, we apply this stylized simulation to explore benchmark holdout design. We find that moving from public benchmarks to private holdout benchmarks shrinks the gap between benchmark scores and user satisfaction on most benchmarks but widens it on a few, depending on where holdout weights shift scoring credit. We stress-test our findings at both the instrument and case-study level, drawing on the V&V framework of Sargent (2013) and GABM-specific evidence criteria. Beyond holdout design, our simulation is a hypothesis-generating sandbox for studying how evaluator and policy choices, in turn, reshape the ecosystem.
Comments70 pages, 15 figures, 23 tables