发表机构
Axya Inc.; McGill University(Axya公司; 麦吉尔大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过360个控制实验,在噪声对话场景下比较多种池化策略,发现注意力池化显著优于其他方法,为工业对话意图分类提供了实证指导。
AI 中文摘要
对话数据库接口面临一个关键挑战:用户自然地将查询嵌入到对话噪声中(问候语、礼貌用语、离题评论),这会降低意图分类的准确性并浪费计算资源。尽管在编排和检索策略方面取得了进展,但一个基本问题仍未得到解答:在生产语言模型中,面对现实对话噪声,哪种池化策略能最大化意图分类准确性?本研究通过360个受控实验弥补了这一空白,这些实验涵盖四种池化配置(均值、最大值、最后标记、注意力以及FFT增强变体),使用Llama-3.2-1B-Instruct在BANKING77和CLINC150数据集上,在干净/噪声条件下使用十个随机种子进行。关键发现表明,注意力池化在噪声条件下始终优于替代策略(相对于默认值约+2.6-2.8 F1),而均值池化性能下降高达约5个F1点。频域滤波未产生一致的准确性提升,主要作为结构变化而非准确性增强组件。这些结果为构建噪声鲁棒的对话分类器提供了具体的、基于证据的指导:对于噪声接口推荐使用注意力池化,应避免均值池化,而最后标记池化适用于干净查询场景。
英文摘要
Conversational database interfaces face a critical challenge: users naturally embed queries in conversational noise (greetings, politeness, off-topic remarks), which degrades intent classification accuracy and wastes computational resources. Despite advances in orchestration and retrieval strategies, a fundamental question remains unanswered: which pooling strategy maximizes intent classification accuracy under realistic conversational noise in production language models? This work addresses this gap through 360 controlled experiments spanning four pooling configurations (mean, max, last-token, attention, and FFT-augmented variants) using Llama-3.2-1B-Instruct on BANKING77 and CLINC150 datasets under clean/noisy conditions with ten random seeds. Key findings reveal that attention pooling consistently outperforms alternative strategies under noisy conditions (~+2.6-2.8 F1 over the default), while mean pooling degrades performance by up to ~5 F1 points. Frequency-domain filtering does not produce consistent accuracy improvements and functions primarily as a structural variation rather than an accuracy-enhancing component. These results provide concrete, evidence-based guidance for building noise-robust conversational classifiers: attention pooling is recommended for noisy interfaces, mean pooling should be avoided, and last-token pooling is appropriate for clean-query scenarios.
Comments11 pages, 4 figures