arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.06664cs.IRcs.CLcs.HC

EviMap:基于证据的分层主题图谱用于探索未标注语料库

EviMap: Evidence-Grounded Hierarchical Topic Maps for Exploring Unlabeled Corpora

Zhiyin Tan, Changxu Duan

首次发表
浏览论文内容

中文总结 AI 辅助

EviMap通过提取文档内证据短语并组织成分层主题图谱,结合LLM和嵌入聚类,为未标注语料库提供可审计的主题概览,确保每个标签可追溯到原始文本。

中文摘要 AI 辅助

研究团队和组织在标签、查询或编码方案存在之前,常常需要探索不熟悉的自由文本集合,包括调查评论、报告和领域文档。在此阶段,第一张主题图谱塑造了用户注意、优先考虑并带入下游分析的内容,因此其可信度应仅限于可验证的程度。现有选项迫使在规模与可验证性之间进行权衡。定性编码保留了证据但速度慢。搜索预设了查询。聚类和主题模型可扩展但产生的标签需要用户解释。一次性大型语言模型(LLM)摘要流畅但难以复现或审计。我们提出了EviMap,一个交互式系统,为研究人员和实践者提供此类语料库的可审计主题概览。在描述语料库和假设的利益相关者关切的模型生成上下文的指导下,EviMap提取文档内的证据短语,并将其(而非整篇文档)组织成方面、组和细粒度主题的三级图谱。基于嵌入的聚类缩小了LLM进行更精细语义判断的搜索空间。每个节点可追溯到支持性短语片段,因此文档通过其包含的证据链接到主题,用户可以根据原始文本审计标签。用户可以从顶层语料库图谱开始,深入主题,检查原始文档中高亮显示的证据,并组合两个主题以查找同时讨论两者的文档。我们在六个异构语料库上展示了这一工作流程,文档数量从2,108到101,699不等,并与平面和分层LLM基线进行了比较。通过将每个标签基于逐字源片段,EviMap使主题图谱不仅可读,而且可验证。代码、演示视频和交互式仪表板可在https URL获取。

英文摘要

Research teams and organizations often explore unfamiliar free-text collections, from survey comments and reviews to reports and domain documents, before labels, queries or coding schemes exist. At this stage, the first thematic map shapes what users notice, prioritize and carry into downstream analysis, so it should be trusted only insofar as it can be verified. Existing options force a trade-off between scale and verifiability. Qualitative coding preserves evidence but is slow. Search presupposes a query. Clustering and topic models scale but produce labels users must interpret. One-shot large language model (LLM) summaries are fluent yet difficult to reproduce or audit. We present EviMap, an interactive system providing researchers and practitioners with an auditable thematic overview of such corpora. Guided by model-generated context describing the corpus and hypothesized stakeholder concerns, EviMap extracts within-document evidence phrases and organizes them, rather than whole documents, into a three-level map of aspects, groups and fine-grained topics. Embedding-based clustering narrows the search space for finer semantic judgments by the LLM. Each node traces back to supporting phrase spans, so documents link to topics through evidence they contain and users can audit labels against the original text. Users can start from a top-level corpus map, drill into topics, inspect highlighted evidence in original documents, and combine two topics to find documents discussing both. We demonstrate this workflow across six heterogeneous corpora spanning 2,108 to 101,699 documents, with a comparison against flat and hierarchical LLM baselines. By grounding every label in verbatim source spans, EviMap makes a topic map not just readable, but verifiable. Code, demo video, and interactive dashboard are available at https://github.com/zhiyintan/EviMap.

发表机构

  • L3S Research Center, Leibniz University Hannover(汉诺威莱布尼茨大学L3S研究中心)
  • Technische Universität Darmstadt(达姆施塔特工业大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑