发表机构
Amazon(亚马逊)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对LLM在文本分类中区分相似标签的难题,提出混淆感知检索与知识注入框架,无需微调即可生成可迁移规则,在多个基准上显著提升Macro F1。
AI 中文摘要
大型语言模型(LLMs)在将文本分类到包含大量语义相似标签的分类体系时存在困难,因为这些差异具有领域特异性,无法通过预训练捕获。为处理大规模标签空间,一种常见方法是通过嵌入相似度检索前K个候选标签,并提示LLM从中选择。然而,前K个检索减少了候选数量,但无法帮助模型区分相似标签。当两个相似标签同时作为候选出现时,模型缺乏在二者间做出正确选择的信号。我们提出一个框架:(1)识别模型难以区分的标签对;(2)扩展候选集以包含易混淆标签;(3)生成针对性规则以区分相似候选。该框架无需微调,且生成的规则可迁移至更小、更经济的模型。在三个基准(WOS、Flipkart、LEDGAR)上,我们的方法相比检索基线将Macro F1提升了最多10.0个百分点,其中小型模型(2B至20B)通过跨模型迁移获得了最多11.5个百分点的提升。
英文摘要
Large language models (LLMs) struggle to classify text into taxonomies with many semantically similar labels, as the distinctions are domain-specific and not captured by pre-training. To handle large label spaces, a common approach retrieves top-$K$ candidate labels by embedding similarity and prompt the LLM to choose among them. However, top-$K$ retrieval reduces the number of candidates but does not help the model tell similar ones apart. When two similar labels both appear as candidates, the model lacks the signal to choose correctly between them. We propose a framework that (1) identifies which label pairs the model struggles to distinguish, (2) expands the candidate set to include confusable labels, and (3) generates targeted rules to differentiate between similar candidates. The framework requires no fine-tuning, and the generated rules transfer to smaller, cheaper models. On three benchmarks (WOS, Flipkart, LEDGAR), our approach improves Macro F1 by up to 10.0pp over retrieval baselines, with smaller models (2B--20B) gaining up to 11.5pp via cross-model transfer.
CommentsEMNLP 2026 (Industry Track)