arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.07093cs.HC

UncertaintyVis:在自动文本到图表生成中保留语言不确定性

UncertaintyVis: Preserving Linguistic Uncertainty in Automated Text-to-Chart Generation

Songheng Zhang, Emily Aurelia, Anthony Tang

首次发表
浏览论文内容

中文总结 AI 辅助

针对自动文本到图表生成会丢弃语言不确定性标记导致读者误判的问题,提出UncertaintyVis系统,通过四类不确定性对应的图表视觉编码实现不确定性保留,实验显示其提升了图表匹配准确率并降低认知需求。

中文摘要 AI 辅助

富含数据的文档将叙事文本与定量主张配对,作者通常会用“几乎”“大约”或“至少”等语言不确定性标记来限定这些主张。自动文本到图表系统会丢弃这些标记,生成的可视化结果看似确定,即便源文本表达的是不确定或不完整的知识,读者可能会过度解读精度并误判作者意图。我们提出UncertaintyVis,这一系统可在自动图表生成过程中保留语言不确定性。对12份文档、8个领域中211个不确定性表达的初步语料分析,得出四类分类:表面形式标准化、精度边界、推理推导和不可推理差距。我们将每个类别映射到特定于图表的视觉编码,这些编码在不破坏读者依赖的空间完整性的情况下传递不确定性,并实现了将大语言模型文本分析与不确定性感知渲染配对的端到端流水线。在由12名参与者组成的两部分研究中,读者将图表与源文本匹配的准确率为85%,将文本与图表匹配的准确率为76%。不确定性感知可视化呈现出认知需求降低的趋势(心理需求和努力的效应量分别为0.460和0.769),75%的参与者更偏好它们而非普通文本,称明确的不确定性编码是验证数据主张的基础。编码效果因图表类型而异:条形图和饼图编码表现稳定,而折线图编码需要重新设计。

英文摘要

Data-rich documents pair narrative text with quantitative claims, and authors routinely qualify those claims with linguistic uncertainty markers such as "nearly," "approximately," or "at least." Automated text-to-chart systems discard these markers, producing visualizations that appear definitive even when the source text expresses hedged or incomplete knowledge. Readers may then over-interpret precision and misjudge author intent. We present UncertaintyVis, a system that preserves linguistic uncertainty during automated chart generation. A formative corpus analysis of 211 uncertainty expressions across 12 documents and 8 domains yielded a four-category taxonomy: Surface Form Normalization, Precision Boundaries, Inferential Derivation, and Non-Inferable Gaps. We mapped each category to chart-specific visual encodings that signal uncertainty without disturbing the spatial integrity readers rely on, and implemented an end-to-end pipeline pairing large language model text analysis with uncertainty-aware rendering. In a two-part study with 12 participants, readers matched charts to source text with 85% accuracy and text to charts with 76%. Uncertainty-aware visualizations trended toward lower cognitive demand (effect sizes 0.460 and 0.769 for mental demand and effort), and 75% of participants preferred them to plain text, describing explicit uncertainty encodings as a basis for verifying data claims. Encoding effectiveness varied by chart type: bar and pie encodings performed consistently, while line chart encodings require redesign.

↑