发表机构
City University of Hong Kong(香港城市大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究比较四种方法在中文隐喻识别中的跨数据集表现,发现专家指导型技能提示可提升跨数据集稳定性,微调则在原生数据集上更具优势。
AI 中文摘要
隐喻识别性能会因文本分布和标注策略不同的数据集而发生显著变化。本研究探究固定的专家指导型流程是否比特定任务的参数适配能产生更均衡的跨数据集表现,针对中文句子级隐喻识别,比较了四种预设条件:BERT微调(BERT-FT)、基于QLoRA的大语言模型微调(LLM-FT)、直接零样本大语言模型提示(LLM-ZS)以及带冻结程序性技能的零样本提示(Skill-ZS)。该技能将涉及语境义、基本义、对比和比较的既定标准进行了操作化定义。评估涵盖CMRE测试集及两个外部数据集CCIME和CMC。微调分数是三个随机种子的均值,而每个零样本分数来自一个确定性配置。微调在原生测试集上仍表现最强:BERT-FT达到91.76的宏F1值。LLM-FT的外部均值最高(83.52),而Skill-ZS接近该值,为82.92,且拥有最高的外部下限(82.64)以及所有三个数据集中最小的观测范围(4.08个百分点)。在匹配的零样本比较中,添加该技能会在每个数据集上减少隐喻预测,这显著降低了CCIME的假阳性,但增加了CMRE测试集和CMC的假阴性。结果表明,专家指导型技能提示是实现更均衡观测到的跨数据集表现的补充途径,而微调在原生数据准确性上仍保持优势。据作者所知,这是首个在中文句子级隐喻识别的相同跨数据集评估中,将专家指导型程序性技能与特定任务微调进行比较的研究。
英文摘要
Metaphor-identification performance can change markedly across datasets that differ in text distribution and annotation policy. We examine whether a fixed expert-informed procedure produces a more even cross-dataset profile than task-specific parameter adaptation. Four prespecified conditions are compared for Chinese sentence-level metaphor identification: BERT fine-tuning (BERT-FT), QLoRA-based large language model fine-tuning (LLM-FT), direct zero-shot LLM prompting (LLM-ZS), and zero-shot prompting with a frozen procedural Skill (Skill-ZS). The Skill operationalizes established criteria involving contextual meaning, basic meaning, contrast, and comparison. Evaluation covers CMRE Test and two external datasets, CCIME and CMC. Fine-tuned scores are means over three seeds, whereas each zero-shot score comes from one deterministic configuration. Fine-tuning remains strongest on the native test set: BERT-FT reaches 91.76 Macro-F1. LLM-FT has the highest external mean (83.52), while Skill-ZS is close at 82.92 and has both the highest external floor (82.64) and the smallest observed range across all three datasets (4.08 points). In the matched zero-shot comparison, adding the Skill reduces metaphorical predictions on every dataset. This sharply lowers false positives on CCIME but increases false negatives on CMRE Test and CMC. The results position expert-informed Skill prompting as a complementary route to more even observed cross-dataset performance, while fine-tuning retains its advantage in native-data accuracy. To our knowledge, this is the first study to compare an expert-informed procedural Skill with task-specific fine-tuning in the same cross-dataset evaluation of Chinese sentence-level metaphor identification.
Comments6 pages, 1 figure, 5 tables