arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.15182cs.AI

VisInteract:面向不完美查询的动态交互式文本到可视化

VisInteract: Towards Dynamic Interactive Text-to-Visualization under Imperfect Queries

  • The Hong Kong Polytechnic University(香港理工大学)
  • ByteDance(字节跳动)

机构由 AI 辅助整理,请以论文原文为准。

Wenxin Xu, Jinwei Lu, Hwanhee Kim, Chen Jason Zhang, Xiao-Yong Wei, Haoyang Li, Yuanfeng Song

AI总结:

针对不完美查询下的文本到可视化问题,提出VisInteract范式与Vis-MCTS算法,通过交互式意图恢复和增强的蒙特卡洛树搜索,显著提升任务成功率。

AI中文摘要:

现实世界中的可视化请求通常存在歧义、不完整或事实错误,然而现有的文本到可视化(Text-to-Vis)系统假设输入是明确指定的,并一次性生成图表。当查询不完美时,系统必须与用户交互以恢复真实意图,但目前没有基准或方法支持这一动态过程。我们引入了VisInteract,一种将文本到可视化重新定义为交互驱动的意图恢复的新范式,以及VisInteract-Bench,据我们所知,这是首个用于动态交互式文本到可视化的基准,具有受控的不完美注入、用于真实多轮反馈的泄漏控制用户代理,以及双视角(代码和图表)的自动评估。在算法方面,我们提出了Vis-MCTS,一种蒙特卡洛树搜索(MCTS)增强方法,对经典MCTS进行了改进,包括渐进式扩展以驯服树搜索中无界的工具参数空间、跨回滚信息共享使澄清和批评惠及整个搜索树,以及维度感知奖励分解,将标量用户反馈沿数据保真度、视觉设计和意图对齐维度进行路由,以解决异构动作间的信用分配问题。在两个LLM骨干网络上的大量实验表明,Vis-MCTS始终优于所有文本到可视化基线,将端到端任务成功率比最强交互式基线提高了13.40%至16.27%,比非交互式基线提高了5倍以上。

英文摘要:

Real-world visualization requests are routinely ambiguous, incomplete, or factually incorrect, yet existing Text-to-Visualization (Text-to-Vis) systems assume well-specified inputs and produce charts in a single pass. When queries are imperfect, a system must \emph{interact} with the user to recover the true intent, but no benchmark or method supports this dynamic process. We introduce \textbf{VisInteract}, a new paradigm that reframes Text-to-Vis as interaction-driven intent recovery, and \textbf{VisInteract-Bench}, to our knowledge, that is the first benchmark for dynamic interactive Text-to-Vis, featuring controlled imperfection injection, a leakage-controlled User Agent for realistic multi-turn feedback, and dual-perspective (code and chart) automated evaluation. On the algorithmic side, we propose \textbf{Vis-MCTS}, a Monte Carlo Tree Search (MCTS) enhanced method, introducing improvements over classical MCTS, that \emph{Progressive Widening} to tame the unbounded tool-argument space in tree search, \emph{cross-rollout information sharing} so clarifications and critiques benefit the entire search tree, and \emph{Dimension-Aware Reward Decomposition} that routes scalar user feedback along data-fidelity, visual-design, and intent-alignment dimensions to resolve credit assignment across heterogeneous actions. Extensive Experiments across two LLM backbones show that Vis-MCTS consistently outperforms all Text-to-Vis baselines, improving end-to-end task success by $13.40\%$--$16.27\%$ over the strongest interactive baseline and by more than $5\times$ over non-interactive ones.

↑