arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.11616cs.AIcs.CVcs.LG

MBA:面向现实世界商业创意的多模态基准与智能体

MBA: Multimodal Benchmark and Agents for Real-World Business Ideation

Hojun Choi, Jaeyo Shin, Suin Lee, Hyunjung Shim

首次发表
浏览论文内容

中文总结 AI 辅助

该研究推出首个多模态商业创意基准MBA-Bench,提出MBA-b和MBA-k两种智能体,经实验其性能显著优于相关基准,为多模态商业创意智能体研究提供了重要支撑。

中文摘要 AI 辅助

由大型语言模型(LLM)驱动的智能体系统为商业创意开发带来了新机遇。然而,现有方法仍局限于纯文本范式,未考虑现实场景固有的多模态特性。为此,我们推出MBA-Bench,这是首个用于训练和评估商业创意智能体的多模态基准,包含六个领域的3万个样本,每个领域具有仅靠文本无法充分传达的独特视觉线索。具体而言,我们自动为图像添加标题,并通过检索查询生成、市场证据检索和证据增强合成,使用GPT-4o为三个商业问题各生成五个参考创意。遵循先前工作,我们使用多模态大语言模型(MLLM)作为评判者,依据六项面向商业的标准评估智能体。为考虑标准隐藏或公开的设置,我们分别提出用于盲设的MBA-b和用于已知设的MBA-k。我们使用两个新的奖励目标——创造性和可行性——训练两者,而MBA-k还针对八项总目标优化六项公开标准。两者均通过基于LoRA的监督微调,再结合针对这些特定设置奖励的分组相对策略优化进行训练。为在MBA-Bench上开展广泛实验,我们设置了两个基准,分别仅适配标题或多模态输入,其中后者在多项指标上接近闭源性能。MBA-b和MBA-k分别比标题基准高出63.9%和77.1%,比多模态基准高出25.6%和35.8%。

英文摘要

Agentic systems powered by large language models (LLMs) have opened new opportunities for business ideation. Yet existing approaches remain confined to a text-only paradigm, despite the inherently multimodal nature of real-world contexts. We thus introduce MBA-Bench, the first multimodal benchmark for training and evaluating business ideation agents, comprising 30K samples across six domains, each domain characterized by distinct visual cues not fully conveyed by text alone. Concretely, we automatically caption images and employ GPT-4o to generate five reference ideas for each of three business questions through retrieval query generation, market evidence retrieval, and evidence-augmented synthesis. Following prior work, we evaluate agents across six business-oriented criteria using MLLM-as-a-Judge. To consider settings where criteria are hidden or disclosed, we present MBA-b and MBA-k for blind and known, respectively. We train both with two novel reward objectives---creativity and feasibility---while MBA-k further optimizes the six disclosed criteria for eight in total. Both are trained via LoRA-based supervised fine-tuning followed by group relative policy optimization with these setting-specific rewards. For extensive experiments on MBA-Bench, we set up two baselines accommodating either captions only or multimodal inputs, with the latter nearing closed-source performance on several metrics. MBA-b and MBA-k outperform caption baselines by 63.9% and 77.1%, and multimodal baselines by 25.6% and 35.8%, respectively.

补充信息

↑