arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2404.01940cs.CL

更好地理解网络犯罪:微调大语言模型在翻译中的作用

Towards Better Understanding of Cybercrime: The Role of Fine-Tuned LLMs in Translation

  • Czech Technical University(捷克技术大学)
  • Rapid7(Rapid7公司)
  • National University of Cuyo (UNCuyo)(库约国立大学)

机构由 AI 辅助整理,请以论文原文为准。

Veronica Valeros, Anna Širokova, Carlos Catania, Sebastian Garcia

更新

中文总结 AI 辅助

针对网络犯罪通信翻译的现有痛点,研究提出用微调大语言模型进行翻译,在NoName057(16)组织的公开聊天数据上验证了其准确性与效率,且成本较人工翻译降低430至23000倍。

中文摘要 AI 辅助

理解网络犯罪通信对网络安全防御至关重要。这通常涉及将通信内容翻译成英文,以便处理、解读并生成及时情报。问题在于翻译难度很大:人工翻译速度慢、成本高且译员稀缺,机器翻译则不够准确且存在偏差。我们提出使用微调的大语言模型(Large Language Models,LLM)生成翻译,以准确捕捉网络犯罪语言的细微差别。我们将该技术应用于俄语黑客活动组织NoName057(16)的公开聊天记录。结果表明,我们的微调LLM模型更优、更快、更准确,且能够捕捉语言的细微差别。该方法证明,实现高保真翻译是可行的,且与人工翻译相比,成本可大幅降低430至23000倍。

英文摘要

Understanding cybercrime communications is paramount for cybersecurity defence. This often involves translating communications into English for processing, interpreting, and generating timely intelligence. The problem is that translation is hard. Human translation is slow, expensive, and scarce. Machine translation is inaccurate and biased. We propose using fine-tuned Large Language Models (LLM) to generate translations that can accurately capture the nuances of cybercrime language. We apply our technique to public chats from the NoName057(16) Russian-speaking hacktivist group. Our results show that our fine-tuned LLM model is better, faster, more accurate, and able to capture nuances of the language. Our method shows it is possible to achieve high-fidelity translations and significantly reduce costs by a factor ranging from 430 to 23,000 compared to a human translator.

补充信息

↑