arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MELLON——面向在线导航的多模态增强型大语言模型

MELLON - Multimodal Enhanced LLM for Online Navigation

Ruiyu Li, Haoyang Cai, Zhitong Guo, Tong Hu

arXiv 2608.09121首次发表:更新:

AI 中文总结

针对网页导航智能体现有模型的不足,本文以WebShop为基准,提出MELLON等三项多模态增强方案,训练1个epoch后MELLON任务完成准确率提升9.26%,验证了多模态方法的重要性。

AI 中文摘要

网页导航智能体能够在不同网站上完成各类任务,当前网页导航领域的基线模型要么是单模态的,要么在处理多模态输入时缺乏强推理能力。针对真实网站模拟环境WebShop基准,本文探索文本与图像的对齐方式,以及多模态推理与规划能力,以提升网页导航智能体的性能。我们提出三项创新的多模态增强方案:面向在线导航的多模态增强型大语言模型(MELLON)、VQAgent和多模态排序器。MELLON在任务完成准确率上实现显著提升,仅训练1个epoch后准确率就提高了9.26%。研究结果表明,有必要进一步探索多模态方法,重点开展更广泛的训练与对齐策略研究,以增强网页导航智能体的有效性。

英文摘要

Web navigation agents are capable of addressing various types of tasks on different websites. Current baselines on web navigation are either unimodal or lack strong reasoning abilities given multimodal inputs. Focusing on the WebShop benchmark, a real-world website simulation, we explore the alignment of text and images, as well as multimodal reasoning and planning abilities, to enhance the performance of web navigation agents. We propose three innovative multimodal enhancements: Multimodal Enhanced LLM for Online Navigation (MELLON), VQAgent, and Multimodal Ranker. MELLON demonstrates a significant improvement in task completion accuracy, with a 9.26% increase after just one epoch of training. Our findings suggest the necessity of further exploration into multimodal approaches, with a focus on more extensive training and alignment strategies to enhance the effectiveness of web navigation agents.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑