发表机构
Universitat Politècnica de Catalunya(加泰罗尼亚理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究实证评估大语言模型在严重类别不平衡下对移动应用评论进行细粒度多标签情感分类,发现少样本提示性能最优(宏F1 0.642),而编码器微调结合生成式数据增强和损失重加权可在更低延迟下接近此性能,为需求工程提供可行方案。
AI 中文摘要
背景:移动应用评论的细粒度情感分类使得需求工程活动能够超越基于极性的观点挖掘,包括基于情感的问题优先级排序和面向特性的反馈分析。然而,从应用评论中自动提取细粒度情感的研究仍然不足。目标:基于先前发布的标注框架和改编自Plutchik分类法的人工标注真实数据,本文研究如何利用大语言模型在严重类别不平衡条件下进行自动多标签情感分类。方法:我们比较了编码器-仅微调在多标签和二值集成公式下的表现,解码器-仅零样本和少样本提示在开源和专有模型上的表现,以及一系列不平衡缓解策略(损失重加权、重采样、生成式数据增强),其中合成评论生成器和提示策略通过内在增强效用排名进行选择。结果:微调编码器在基线(多标签:0.387;二值集成:0.450)上大幅落后于最佳解码器少样本提示(宏F1 0.642);将最佳多标签编码器与生成式数据增强和正加权损失相结合,缩小了大部分差距(+0.204),且推理延迟比解码器低至三个数量级,在最稀有情感上增益最大,从未被检测到提升至最高+0.501 F1。结论:大语言模型使得应用评论的细粒度多标签情感分类对需求工程流程可行,宏F1适中,最佳公式和缓解策略依赖于骨干模型和公式。我们发布了实验流程、合成语料库和微调检查点,以供复现和重用。
英文摘要
Context: Fine-grained emotion classification of mobile app reviews enables requirements engineering activities that go beyond polarity-based opinion mining, including emotionally informed issue prioritisation and feature-oriented feedback analysis. However, automatic fine-grained emotion extraction from app reviews remains understudied. Objectives: Building on a previously published annotation framework and human-labelled ground truth adapted from Plutchik's taxonomy, this paper investigates how large language models can be leveraged for automatic multi-label emotion classification under severe class imbalance. Methods: We compare encoder-only fine-tuning under multi-label and binary-ensemble formulations, decoder-only zero- and few-shot prompting across open-source and proprietary models, and a catalogue of imbalance mitigation strategies (loss reweighting, resampling, generative data augmentation), with the synthetic-review generator and prompting strategy selected via an intrinsic augmentation-utility ranking. Results: Fine-tuned encoders trail the best decoder-only few-shot prompting (macro-F1 0.642) by a wide margin at baseline (multi-label: 0.387; binary ensemble: 0.450); pairing the best multi-label encoder with generative data augmentation and positive-weighted loss closes most of this gap (+0.204) at up to three orders of magnitude lower inference latency than the decoders, with the largest gains on the rarest emotions, from undetected to gains of up to +0.501 F1. Conclusion: Large language models make fine-grained, multi-label emotion classification of app reviews feasible for requirements engineering pipelines, with modest macro-F1, and the best formulation and mitigation strategy are backbone- and formulation-dependent. We release the experimental pipeline, synthetic corpora, and fine-tuned checkpoints for replication and reuse.