发表机构
Tableau Research(Tableau研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对对话式可视化分析智能体评估难题,提出无参考指标Lexara-RF,仅用提示、数据和响应评分,通过13个指标实现验证式评估,性能与参考法相当且优于基线。
AI 中文摘要
由大语言模型驱动的对话式可视化分析(CVA)智能体能够根据开放式查询生成可视化和自然语言解释。评估这些多模态输出具有挑战性:精心策划的参考基准编写成本高昂,无法全面覆盖有效响应的空间,并且在生产环境中不可用。基于Lexara评估框架,我们引入了Lexara-RF,这是一组无参考指标,仅使用提示、数据和模型响应来对CVA输出进行评分。我们将评估重新定义为验证:13个指标将可视化设计理论和格赖斯合作原则具体化为可计算的一致性、意图对齐和设计有效性检查。在一个人工评分的CVA测试用例语料库上,Lexara-RF实现了与基于参考的公式相当的对齐,优于表面相似性的自然语言生成基线,并以高精度定位了结构上扎根的失败。
英文摘要
Conversational visual analytics (CVA) agents powered by large language models generate visualizations and natural-language explanations from open-ended queries. Evaluating these multimodal outputs is challenging: curated reference benchmarks are costly to author, cannot comprehensively capture the space of valid responses, and are unavailable in production. Building on the Lexara evaluation framework, we introduce Lexara-RF, a reference-free set of metrics that scores CVA outputs using only the prompt, data, and model response. We reformulate evaluation as verification: 13 metrics operationalize visualization design theory and Gricean cooperative principles as computable consistency, intent-alignment, and design validity checks. On a human-rated corpus of CVA test-cases, Lexara-RF achieves alignment comparable to reference-based formulations, outperforms surface-similarity NLG baselines, and localizes structurally grounded failures with high accuracy.
Comments4 pages, 1 figure Conversational Visual Analytics, Evaluation Metrics, Visualization Design, Cooperative Communication Principles
Journal ref2026 IEEE Visualization and Visual Analytics (VIS) Conference