arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.18437cs.IR

探索LLMs和RAG在车辆零部件合理且可解释材料预测中的应用

Exploring LLMs and RAG for Plausible and Explainable Material Prediction of Vehicle Components

  • University of Stuttgart(斯图加特大学)

机构由 AI 辅助整理,请以论文原文为准。

Frederik Wagner, Annerose Eichel, Sabine Schulte im Walde

中文总结 AI 辅助

本研究探索LLM及RAG在车辆零部件材料预测中的可解释性,发现生成式LLM显著优于先前工作,RAG未进一步超越,并指出超参数、语料库和评估设计等挑战。

中文摘要 AI 辅助

在这项工作中,我们探讨了大型语言模型(LLMs)是否能够在无需大量微调的情况下,准确预测并解释车辆零部件(如刹车盘或燃油喷射器)的合理材料。我们测试并评估了三种方法:标准生成式LLM基线、单次检索增强生成(RAG)方法以及迭代式链式验证(CoVe)变体。对于检索,我们依赖公开可用的数据,使用领域过滤的维基百科语料库。由于该任务不存在黄金标准,我们开发了一个自定义的基于网络的注释工具,支持结构化领域专家评估的关键功能。基于LLM的生成性能显著优于先前工作,而所测试的RAG方法并未进一步超越该性能。我们的结果揭示了基于RAG系统面临的剩余挑战:超参数优化、高质量且法律上可获取的领域语料库的可用性,以及专家评估研究设计。

英文摘要

In this work, we explore whether LLMs can accurately predict and explain plausible materials for vehicle components such as brake discs or fuel injectors without requiring extensive fine-tuning. We test and evaluate three approaches: a standard generative LLM baseline, a single-pass Retrieval-Augmented Generation (RAG) approach, and an iterative Chain-of-Verification (CoVe) variant. For retrieval, we rely on publicly available data using a domain-filtered Wikipedia corpus. Since no gold standard exists for this task, we develop a custom web-based annotation tool supporting crucial functions for structured domain expert evaluation. LLM-based generation substantially outperforms prior work, which is not further surpassed by the tested RAG approaches. Our results surface remaining challenges for RAG-based systems: hyperparameter optimization, the availability of high-quality, legally accessible domain corpora, and expert evaluation study design.

↑