arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

大语言模型在可视化领域特定语言(DSL)上的失败之处

Where LLMs Fail with Visualization DSLs

Chang Han, Andrew McNutt, Katherine Isaacs

arXiv 2610.01873首次发表:更新:

AI 中文总结

本研究评估了3个LLM在10种JSON风格可视化DSL上的41项任务,识别出四种反复出现的失败模式,并将其与特定DSL特征关联,为未来DSL设计提供指导。

AI 中文摘要

随着大语言模型(LLM)承担起使用可视化领域特定语言(DSL)编写图表的角色,塑造这些语言的人类约束可能不再适用,因为对人类来说容易的事情对模型来说未必容易。为了理解LLM如何更好地与DSL协作,我们探索了它们在当前DSL设计中的失败位置和方式。我们使用3个LLM对10种JSON风格的可视化DSL进行了41项任务的评估,然后通过JSON和渲染检查以及失败案例的定性编码来评估生成的规范。通过分析这种规范生成过程的失败情况,我们识别出四种反复出现的失败模式,将每种模式与特定的DSL特征联系起来,并讨论了未来DSL设计的设计考量。

英文摘要

As LLMs take up the role of authoring charts using visualization domain-specific languages (DSLs), the human constraints that shaped those languages may no longer apply, as what is easy for a person is not necessarily easy for a model. To understand how LLMs might work better with DSLs, we explore where and how they fail with current DSL designs. We evaluate 10 JSON-style visualization DSLs with 41 tasks across 3 LLMs, then assess the generated specifications with JSON and rendering checks, and qualitative coding of failed cases. Analyzing how this specification generation process fails, we identify four recurring failure patterns, link each to specific DSL features, and discuss design considerations for future DSL designs.

CommentsVIS 2026 VISxGenAI, 6 pages, 3 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑