arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DashArena:面向交互式分析仪表板生成的大语言模型基准测试

DashArena: Benchmarking LLMs on Interactive Analytic Dashboard Generation

Xiaotong Wang, Dazhen Deng

arXiv 2608.10567首次发表:更新:

AI 中文总结

研究人员推出首个交互式分析仪表板生成基准DashArena,含轨迹回放与VLM评判,提炼出DashJudge-8B,发现前沿模型仍存多类缺陷,交互评估可捕捉遗漏问题。

AI 中文摘要

分析仪表板整合了协同视图与交互功能,用于数据探索和决策制定。现有模型可根据数据和自然语言目标生成分析仪表板,但评估其实用性仍存在困难。仪表板生成属于开放式任务,仅通过静态外观或成功执行无法全面衡量其分析支持能力与交互质量。我们推出DashArena,据我们所知,这是首个针对开放式、任务导向型交互式分析仪表板生成的基准测试。其核心创新在于要求每个系统同时生成仪表板和可复现的交互轨迹:浏览器执行器会回放该轨迹,将系统预设的分析工作流转化为可复现的视觉与执行证据;视觉语言模型(VLM)评判器基于该证据对候选系统进行比较,再通过布拉德利-特里(Bradley-Terry)聚合生成排行榜。我们进一步将该评判器提炼为开放权重模型DashJudge-8B。人工评估显示,DashJudge-8B可有效复现人类评判结果; ablation实验表明,交互证据能提升评判一致性。对前沿模型的实验揭示了其在渲染、分析及交互方面存在的持续缺陷。综上,这些结果表明,现实场景中的仪表板生成仍具挑战性,且感知交互的评估方式能捕捉到仅靠静态或仅执行检查所遗漏的缺陷。

英文摘要

Analytic dashboards combine coordinated views and interactions for data exploration and decision-making. Recent models can generate them from data and natural-language goals, but evaluating their usefulness remains difficult. Dashboard generation is open-ended, and neither static appearance nor successful execution alone captures analytical support and interaction quality. We introduce DashArena, to our knowledge the first benchmark for open-ended, task-grounded generation of interactive analytic dashboards. Its key innovation is to require each system to generate both a dashboard and a replayable interaction trajectory. A browser executor replays the trajectory and turns the system's intended analytical workflow into reproducible visual and execution evidence. A VLM judge compares candidates using this evidence, and Bradley--Terry aggregation produces the leaderboard. We further distill the judge into the open-weight DashJudge-8B. Human evaluations show that DashJudge-8B effectively reproduces human judgments and ablations show that interaction evidence improves judge agreement. Experiments with frontier models reveal persistent rendering, analytical, and interaction failures. Together, these results show that realistic dashboard generation remains challenging and that interaction-aware evaluation captures failures missed by static or execution-only checks.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑