发表机构
Georgia State University; Amazon; Central China Normal University(佐治亚州立大学; 亚马逊; 华中师范大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对大规模多标签文本分类,提出DualMLC双分支框架,融合自回归解码器与双向编码器的异构表示,通过后期逻辑融合提升性能,在三个基准上达到最先进结果。
AI 中文摘要
大规模多标签文本分类是从包含数千或数万个候选标签的词汇表中为每个文档分配一小部分相关标签。尽管预训练语言模型改善了语义文本表示,但大多数基于表示的方法将其预测流程集中在主编码器上,或在单个排序器内组合辅助特征。因此,异构语言模型之间的互补性仍未得到充分探索。我们提出了DualMLC,一种双分支框架,通过自回归仅解码器语言模型和双向编码器处理同一文档。每个分支维护自己的表示路径,并在共享标签空间上独立估计相关性分数。DualMLC通过后期逻辑融合结合两个分数向量,使共享证据能够强化相关标签,而分支特定证据能够补偿另一分支表示的局限性。DualMLC在三个广泛使用的大规模多标签文本分类基准上取得了最先进的结果。消融研究进一步证实,集成异构预测器比单独任一分支产生更强的排序。源代码可在此https URL公开获取。
英文摘要
Large-scale multi-label text classification assigns a small subset of relevant labels to each document from a vocabulary containing thousands or tens of thousands of candidate labels. Although pretrained language models have improved semantic text representations, most representation-based approaches center their prediction pipelines on a primary encoder or combine auxiliary features within a single ranker. The complementarity between heterogeneous language models therefore remains insufficiently explored. We propose DualMLC, a dual-branch framework that processes the same document through an autoregressive decoder-only language model and a bidirectional encoder. Each branch maintains its own representation pathway and independently estimates relevance scores over the shared label space. DualMLC combines the two score vectors through late logit fusion, allowing shared evidence to reinforce relevant labels and branch-specific evidence to compensate for limitations in the other branch's representation. DualMLC achieves state-of-the-art results on three widely used large-scale multi-label text classification benchmarks. Ablation results further confirm that integrating the heterogeneous predictors produces stronger rankings than either branch alone. The source code is publicly available at https://github.com/huiyegit/DualMLC.