预训练模型在驾驶行为视频描述中的适配方法比较研究
Comparative study of adapting pre-trained models for driving behavior video captioning
浏览论文内容
中文总结 AI 辅助
本研究比较了多种微调和提示方法在自动驾驶视频数据集上适配LLM的效果,全微调在BDD-X上表现良好,部分指标超越基线,并探讨了LoRA和提示工程的局限性。
中文摘要 AI 辅助
本报告考察并比较了现有的多种微调和提示方法,并将其应用于自动驾驶领域。其思路是通过在视频数据集上适配大型语言模型(LLM)来比较这些方法。LLM在理解不同形式数据方面已经变得非常出色,本研究旨在将驾驶情境的低维理解引入我们的主要测试模型SpaceTimeGPT。在BDD-X(伯克利深度驾驶解释)数据集上的实验表明,全微调框架在部分自动评估指标上表现良好,甚至在某些指标上超越了基线。我们还尝试了低秩适配(LoRA)和提示工程在VideoLLaVA模型上的应用,并讨论了其局限性。
英文摘要
This report examines and compares some of the many fine tuning and prompting methods existing, applying them within the domain of autonomous driving. The idea is to compare these methods by adapting a Large Language Model (LLM) on a video dataset. LLM's have become extremely good at achieving a good understanding of different forms of data and this study aims to induce a low dimensional understanding of driving situations into our primary test model SpaceTimeGPT. Experiments on BDD-X (Berkeley DeepDrive eXplanation) dataset demonstrate good performance of the full fine tuning framework on some automatic metrics, and in some metrics, it even surpasses the baseline. We also try Low-Rank Adaptation (LoRA) and prompt engineering on VideoLLaVA model and discuss its limitations.
发表机构
- University of Tübingen(蒂宾根大学)
- Bosch Center for Artificial Intelligence(博世人工智能中心)
机构由 AI 辅助整理,请以论文原文为准。