arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.21408cs.AI

罗马乌尔都语仇恨言论分类:参数高效微调与提示工程的对比研究

Hate Speech Classification In Roman Urdu: A Comparative Study On Parameter Efficient Fine-Tuning And Prompt Engineering

Toneema Zubair

首次发表
浏览论文内容

中文总结 AI 辅助

本研究针对低资源的罗马乌尔都语,对比了LLM零样本推理、LoRA的PEFT、混合/人工提示调优、零/少样本提示工程四种仇恨言论分类方法,以确定最有效的技术。

中文摘要 AI 辅助

由于互联网和社交媒体的广泛普及,有毒和仇恨内容呈指数级增长,造成了严重的困扰和负面的社会影响。罗马乌尔都语是一种低资源语言,在巴基斯坦及全球乌尔都语使用者社区中使用,由于其非正式语法、不一致的句子结构以及单词拼写的多种变体,带来了额外的挑战。本研究旨在确定在数据有限的低资源环境中进行仇恨言论分类的最有效技术。为解决这一问题,研究调查并比较了最新方法,包括提示调优、使用LoRA的参数高效微调(PEFT)以及提示工程,在各种实验配置下开展。为实现该目标,设计了四项实验:第一项实验是直接使用大语言模型(LLM)进行推理,不进行任何微调,以评估这些模型在零样本设置下对罗马乌尔都语的理解程度,尤其是在数据有限的情况下;第二项实验是使用LoRA进行参数高效微调(PEFT),仅更新一小部分参数,从而降低计算成本;第三项实验探索了混合提示和人工设计提示的提示调优,使用相对于整个数据集而言非常小的训练示例集,因此也具有计算效率;最后,第四项实验通过零样本和少样本学习应用提示工程,仅依靠精心设计的指令提示进行分类,无需进一步训练。

英文摘要

Due to the widespread accessibility of the internet and social media, toxic and hateful con-tent has grown exponentially, causing significant distress and negative societal impacts. Ro-man Urdu, a low-resource language used in Pakistan and among Urdu-speaking communities worldwide, presents additional challenges because of its informal grammar, inconsistent sen-tence structures, and multiple variations in word spellings. This research aims to identify the most effective techniques for hate speech classification in such low-resource settings with limited data. To address this, the study investigates and compares the latest approaches, in-cluding prompt tuning, parameter-efficient fine-tuning (PEFT) using LoRA, and prompt en-gineering, under various experimental configurations. To achieve this objective, four exper-iments were designed. The first experiment involved direct inferencing with LLMs without any fine-tuning, to evaluate how well these models understand Roman Urdu in a zero-shot setting, especially given limited data. The second experiment utilized parameter-efficient fine-tuning (PEFT) with LoRA, which updates only a small subset of parameters, thereby reducing computational cost. The third experiment explored prompt tuning with both mixed and manually crafted prompts, using very small sets of training examples relative to the entire dataset, making it computationally efficient as well. Finally, the fourth experiment applied prompt engineering through zero-shot and few-shot learning, relying solely on care-fully designed instruction prompts for classification without further training.

↑