AI 中文总结
MatrAIx是含83亿 persona 智能体的模拟用户评估基础设施,含Persona 8B、MatrAIx Playground及1010个多领域任务,经验证可高效评估AI系统与数字产品。
AI 中文摘要
对AI系统和数字产品进行人工评估成本高、速度慢且难以规模化,离线评估虽更具可扩展性,但往往忽略了人类的多样性和交互行为。为此,我们推出了MatrAIx,这是一种用于测试AI系统和数字产品的人口规模模拟用户评估基础设施,适用于具有异质性的用户。MatrAIx包含三个核心组件:第一,Persona 8B包含83亿个 persona 记录,由1290个分类维度表示,这些记录要么从保留相关属性的依赖图中采样,要么源自人工编写的用户画像;我们发布了一个经过质量筛选的核心子集,约100万个 persona,其中包括599847个基于人类的记录和400000个合成记录。第二,MatrAIx Playground提供了四个环境,供不同用户评估和与数字产品交互:调查、AI聊天机器人、网页和应用。第三,MatrAIx提供了1010个应用任务,涵盖超过25个领域,包括商业、软件、金融和医疗保健。我们在8个代表性任务中进行了18189次评估试验,Persona 智能体由三个大语言模型(LLM)提供支持:Claude Opus 4.8、GPT 5.5和Claude Haiku 4.5。所得反馈捕捉了不同 persona 背景下决策和偏好的差异,包括价格上涨后的犹豫、AI助手失败后继续使用的意愿以及延迟容忍度。我们开展了两项主要验证研究:第一项是包含400次试验的对照研究,评估了所有四个环境中10个行为属性下的 persona 依从性,在366次试验(占比91.5%)中,声明的行为得到了表达或正确抑制;第二项是由人工和LLM评判员对基于人类的 persona 的提取质量进行评估。总体而言,MatrAIx为使用多样化的模拟人类用户评估AI系统和数字产品提供了端到端的基础设施。
英文摘要
Human evaluation of AI systems and digital products is costly, slow, and difficult to scale. Offline evaluations are more scalable but often abstract away human diversity and interactive behavior. We therefore introduce MatrAIx, a population-scale simulated-user evaluation infrastructure for testing AI systems and digital products with heterogeneous users. MatrAIx has three core components: First, Persona 8B contains 8.3 billion persona records represented by 1,290 categorical dimensions. Records are either sampled from a dependency graph that preserves correlated attributes or derived from human-authored profiles. We release a quality-filtered coreset of approximately 1 million personas, comprising 599,847 human-grounded and 400,000 synthetic records. Second, the MatrAIx Playground provides four environments in which diverse users evaluate and interact with digital products: Survey, AI Chatbot, Web, and App. Third, MatrAIx provides 1,010 application tasks spanning more than 25 domains, including Commerce, Software, Finance, and Healthcare. We conducted 18,189 evaluation trials across eight representative tasks. Persona agents were powered by three LLMs: Claude Opus 4.8, GPT 5.5, and Claude Haiku 4.5. The resulting feedback captures how decisions and preferences vary across persona backgrounds, including hesitation after a price increase, willingness to continue after an AI assistant fails, and latency tolerance. We conducted two main validation studies: First, a 400-trial controlled study evaluated persona adherence across ten behavioral attributes and all four environments. The declared behavior was expressed or correctly suppressed in 366 trials (91.5%). Second, human and LLM judges evaluated the extraction quality of human-grounded personas. Overall, MatrAIx provides an end-to-end infrastructure for evaluating AI systems and digital products with diverse simulated human users.
CommentsProject website: https://matraix.ai