离散扩散语言模型是无需训练的多标签分类器
Discrete Diffusion Language Models Are Training-Free Multi-Label Classifiers
- International Institute of Information Technology, Hyderabad(印度海得拉巴国际信息技术学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究提出无需训练的多标签分类方法dLLM-SetScore,基于离散掩码扩散语言模型,在多个数据集上对比基准模型,验证了其性能优势并完成了相关理论分析。
AI中文摘要:
我们提出了dLLM-SetScore,这是一种无需训练的方法,它使用离散掩码扩散语言模型进行多标签文本分类。对于每个候选标签,它会提出一个简短的是/否问题,并在一个掩码位置比较两个答案标记的概率。该方法无需进行特定任务的微调,也无需在文本蕴含数据集上进行训练;仅用200个示例的标注验证切片即可选择阈值、温度和提示措辞。我们首先证明,将所有标签放在一个提示中会产生强烈的槽位位置不对称性:在99.4%的GoEmotions示例和100%的Reuters示例中,第一个答案槽位被预测为正类。按标签评分会将每个标签置于相同的句法位置,从而使预测对标签顺序不变,并避免了这种人为偏差。我们在六个数据集上评估了LLaDA-8B和Dream-7B,将其与NLI模型、自回归大语言模型(LLM)、SetFit以及监督分类器进行对比。在两个扩散模型家族共有的五个数据集上,指令调优检查点在10次对比中有9次提升了宏F1值,在10次对比中有8次提升了微F1值,不过这些对比并未确定其原因。在我们的方案中,LLaDA-Instruct在Reuters和ECtHR指标上均达到了最高的无需训练的值。我们证明了置换不变性,描述了加权汉明损失下的阈值决策,并推导了召回率和F1的候选上限。探索性的局部联合集细化步骤会降低有偏和无偏初始值的F1值,因此作为负结果被保留。
英文摘要:
We present dLLM-SetScore, a training-free method that uses discrete masked-diffusion language models for multi-label text classification. For each candidate label, it asks a short yes/no question and compares the probabilities of the two answer tokens at one masked position. The method uses no task-specific fine-tuning or training on textual-entailment datasets; a 200-example labelled validation slice selects thresholds, temperature, and prompt wording. We first show that placing all labels in one prompt creates a strong slot-position asymmetry: the first answer slot is predicted positive on $99.4\%$ of GoEmotions examples and $100\%$ of Reuters examples. Per-label scoring places every label in the same syntactic position, making predictions invariant to label order and avoiding this artifact. We evaluate LLaDA-8B and Dream-7B on six datasets against NLI models, an autoregressive LLM, SetFit, and supervised classifiers. On the five datasets shared by both diffusion families, Instruct checkpoints improve macro-F1 in 9 of 10 comparisons and micro-F1 in 8 of 10, although these comparisons do not identify the cause. Within our protocol, LLaDA-Instruct records the highest training-free values for both Reuters and ECtHR metrics. We prove permutation invariance, characterize thresholded decisions under weighted Hamming loss, and derive shortlist ceilings for recall and F1. An exploratory local Joint Set Refinement step lowers F1 from biased and unbiased initializations and is retained as a negative result.