发表机构
SRI International(SRI国际)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出MIL-BERT算法,基于多实例学习训练神经网络选择文本片段进行分类,可处理近100万token样本,在3个数据集上取得SOTA结果,且弱标记训练模型可泛化至小实例分类。
AI 中文摘要
许多文本分类决策仅基于构成片段即可做出。受多实例学习领域的启发,我们提出了一种训练神经网络以通过选择此类片段对文本进行分类的算法。我们证明该方法具有可扩展性,已在近100万token的样本上验证了学习能力。我们在7个数据集上评估了方法,重点关注远超基础模型编码限制的长文本集合。该算法在3个数据集上取得了SOTA结果:新闻媒体的政治偏见识别、长篇故事的触发警告识别、推文集合中作者的人口统计特征识别。此外,在弱标记文本包(bags)上训练的模型可泛化至准确分类更小的构成实例。除了这些问题上的新SOTA,该方法是少数在这些数据集上表现优异的神经方法之一。
英文摘要
Many text classification decisions are viable based on constituent excerpts alone. Taking inspiration from the field of multiple instance learning, we present an algorithm for training a neural network to classify text by selecting such excerpts. We show that our approach is also scalable with demonstrated learning against samples with nearly 1M tokens. We evaluate our methods on 7 datasets with emphasis on long-textual collections that far exceed the encoding limit of our base model. We present state-of-the-art results with this algorithm on 3 datasets: identification of political bias in news outlets, trigger warnings in long stories, and demographic characteristics of authors in tweet collections. Furthermore, the model trained on weakly-labeled collections of text (bags) generalizes to accurately classify constituent, smaller instances. Besides a new state-of-the-art for these problems, this approach is one of the few neural methods to excel in these datasets.