arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

IRWOZ 2.0:面向工业机器人对话的大语言模型驱动对话数据集

IRWOZ 2.0: A Large Language Model-driven Dialogue Dataset for Industrial Robot Conversations

Chen Li, Dimitrios Chrysostomou

arXiv 2609.04030首次发表:更新:

发表机构

Aalborg University(奥尔堡大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对IRWOZ初始版本对话状态与话语噪声导致状态跟踪准确率受限的问题,提出IRWOZ 2.0数据集,结合Mistral/Claude-3.5的LLM增强生成与质量优化,经实验使GPT-2的BLEU-4得分大幅提升,已公开以支撑工业HRI研究。

AI 中文摘要

IRWOZ通过领域特定注释改进了工业人机交互(HRI)对话系统,但初始版本的对话状态和话语存在大量噪声,限制了状态跟踪准确率。本文介绍IRWOZ 2.0,该版本通过大语言模型(LLM)增强生成(Mistral/Claude-3.5)及质量优化解决了上述局限。改进后的数据集扩展至4个工业领域(装配、配送、定位、重定位)的390段对话,包含人工修正及自动错别字移除。对话状态跟踪的基准实验显示出显著提升,与原始IRWOZ相比,GPT-2的BLEU-4得分从0.1651提升至0.5604。为支持工业HRI研究,我们在该公开网址发布了IRWOZ 2.0数据集。

英文摘要

IRWOZ has improved industrial human-robot interaction (HRI) dialogue systems through domain-specific annotations. However, its initial version contains substantial noise in dialogue states and utterances, limiting state-tracking accuracy. We introduce IRWOZ 2.0, which addresses these limitations through large language model (LLM) enhanced generation (Mistral/Claude-3.5) and quality refinements. Our improved dataset expands to 390 dialogues across 4 industrial domains (Assembly, Delivery, Position, Relocation), featuring manual corrections and automated typo removal. Benchmark experiments on dialogue state tracking demonstrate significant improvements, with GPT-2's BLEU-4 score increasing from 0.1651 to 0.5604 compared to original IRWOZ. To support industrial HRI research, we publicly released IRWOZ 2.0 dataset at https://ieee-dataport.org/documents/irwoz-20-large-language-model-driven-dialogue-dataset-industrial-robot-conversations

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑