AI 中文总结
研究针对预训练语言模型调整难的问题,提出KeySI交互框架,通过基于关键词的概念规范实现特征级反馈,减少人工文档处理需求,经用户研究等评估,证明其能有效捕捉用户意图并改善嵌入对齐。
AI 中文摘要
在大规模文本分析任务中,预训练语言模型常被用于嵌入文本语料库以进行下游分析。但这些模型可能难以捕捉特定领域语义,且调整通常需要大量标注数据和技术专长。近期方法通过文档投影中的视觉交互捕捉人类反馈作为模型调整的训练信号。然而,这些方法基于文档级反馈,需用户打开并评估单个文档才能提供有效反馈。本文提出KeySI,一个通过基于关键词的概念规范实现特征级反馈的交互框架。用户通过将提取的关键词组织成代表概念的组来指定反馈,KeySI将其转化为文档级监督用于后续调整。通过以关键词作为主要交互媒介,KeySI减少了人工文档检查和标注的需求,降低了调整嵌入模型的障碍。我们展示了一个原型实现,给定一个语料库,策划代表性关键词,通过降维可视化关键词和文档嵌入,允许交互式指定关键词组,并通过系统反馈支持迭代优化。我们通过用户研究、使用场景和定量实验评估KeySI,证明其在捕捉用户意图和改善嵌入对齐方面的有效性。
英文摘要
In large-scale text analysis tasks, pre-trained language models are often used to embed text corpora for downstream analysis. However, such models may struggle to capture domain-specific semantics and adapting them typically requires large amounts of labeled data and technical expertise to implement training pipelines. Recent approaches have demonstrated how visual interactions in document projections can capture human feedback as training signals for model tuning. However, these methods operate on document-level feedback, which requires users to open and assess individual documents in order to provide effective feedback. In this paper, we propose KeySI, an interaction framework that enables feature-level feedback through keyword-based concept specification. Users specify feedback by organizing extracted keywords into groups representing concepts, which KeySI translates into document-level supervision for subsequent tuning. By operating on keywords as the primary interaction medium, KeySI reduces the need for manual document inspection and labeling and lowers the barrier to adapting embedding models. We present a prototype implementation that, given a corpus, curates representative keywords, visualizes keywords and document embeddings via dimensionality reduction, allows interactive specification of keyword groups, and supports iterative refinement through system feedback. We evaluate KeySI through a user study, usage scenarios, and quantitative experiments demonstrating its effectiveness in capturing user intent and improving embedding alignment.
CommentsAccepted to IEEE VIS 2026