发表机构
Universitat Ramon Llull(拉蒙·柳利大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出零样本可视化(ZSV)任务,利用大型语言模型将文档映射到用户自然语言指定的概念轴,通过基准评估发现基于下一个词元概率的评分方法在语义保真度与计算成本间取得最佳平衡,并为构建端到端系统提供实用指南。
AI 中文摘要
我们研究了将大型语言模型(LLMs)应用于文本语料库的视觉探索。我们引入了零样本可视化(ZSV)任务,在该任务中,用户以自然语言指定概念,文档被映射到相应的概念轴上以进行可视化。构建具有实用价值的ZSV系统并非易事,因为它需要在特征函数、高效实现权衡以及影响可视化质量的前/后处理决策的交汇处做出选择。为此,我们建立了一个基准,比较了在此设置中涵盖嵌入相似性、直接语义判断和条件似然估计的方法。跨多个数据集和使用案例,我们评估了不同评分方法和设计选择在语义保真度、评分保真度和计算成本方面的特性。我们的结果表明,基于下一个词元概率的评分在评估的方法中提供了最强的实用权衡。我们进一步将此方法应用于未标记的语料库,以检查其在现实探索性设置中的行为。这些实验突出了额外的设计考虑,包括使用分级轴与二元相关性过滤相结合,并揭示了离题文档中的组合情感偏差。基于这些发现,我们为构建端到端的ZSV基线提供了实用指南。
英文摘要
We study the application of large language models (LLMs) to the visual exploration of textual corpora. We introduce zero-shot visualization (ZSV), a task in which users specify concepts in natural language and documents are mapped onto the corresponding concept axes for visualization. Building a ZSV system of practical value is non-trivial, as it requires choices at the intersection of feature functions, efficient implementation tradeoffs, and pre/post-processing decisions affecting visualization quality. To that end, we establish a benchmark that compares methods spanning embedding similarity, direct semantic judgments, and conditional likelihood estimation in this setting. Across multiple datasets and use cases we evaluate the properties of different scoring methods and design choices in terms of semantic faithfulness, score fidelity, and computational cost. Our results identify that scoring based on next-token probabilities offers the strongest practical trade-off among the evaluated methods. We further apply this approach to unlabeled corpora to examine its behavior in realistic exploratory settings. These experiments highlight additional design considerations, including the use of graded axes together with binary relevance filtering, and reveal a compositional sentiment bias in off-topic documents. Based on these findings, we provide practical guidelines for constructing end-to-end ZSV baselines.