发表机构
University of Padova; Universitat Pompeu Fabra(帕多瓦大学; 庞培法布拉大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
PANEL是一个开源、自托管的Web平台,用于对生成模型进行受控人工评估,支持多种刺激类型、复杂调查逻辑和统计分析,并符合GDPR要求。
AI 中文摘要
人工判断是评估生成模型的参考标准,然而用于收集这些判断的软件却落后于方法论的发展。研究人员会调整为感知协议(如MUSHRA)设计的听力测试框架,依赖封闭的商业调查平台,或实现一次性的Web应用程序。诸如Chatbot Arena和Music Arena等实时竞技场对公开部署的系统进行大规模排名,但不支持在实验室自己的参与者中对实验室自身模型进行受控比较。我们提出了PANEL,一个用于此类研究的开源、自托管平台。研究在浏览器中编写,并以单个链接分发,支持音频、视频、图像和文本刺激,七种问题类型,以及筛选和跳过逻辑。该平台报告每个问题的摘要、跨条件的显著性检验、成对胜率和Bradley-Terry分数,并支持基于试点数据的功效分析。同意版本管理、自助撤回、保留执行和审计日志支持符合GDPR的操作。每项研究导出为机器可读的规范。PANEL可在https URL获取。
英文摘要
Human judgement is the reference measure for evaluating generative models, yet the software used to collect it lags behing the methodology. Researchers adapt listening-test frameworks designed for perceptual protocols such as MUSHRA, rely on closed commercial survey platforms, or implement single-use web applications. Live arenas such as Chatbot Arena and Music Arena rank publicly deployed systems at scale, but do not support controlled comparisons of a laboratory's own models with its own participants. We present PANEL, an open-source, self-hosted platform for such studies. A study is authored in the browser and distributed as a single link, with audio, video, image, and text stimuli, seven question types, and screening and skip logic. The platform reports per-question summaries, across-condition significance tests, pairwise win rates and Bradley--Terry scores, and supports power analysis from pilot data. Consent versioning, self-service withdrawal, retention enforcement, and audit logging support GDPR-compliant operation. Each study exports as a machine-readable specification. PANEL is available at https://github.com/matteospanio/panel.
Comments3 pages, 2 figures, ISMIR 2026 Late Breaking Demo