arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MetaEncoder:探索带自然语言接口的多模态System One决策双编码器的极限

MetaEncoder: Exploring the Limit of Bi-Encoders for Multimodal System One Decision Making with Natural Language Interface

Jianpeng Cheng, Guangyu Sun, Aashu Singh, Benyu Zhang, Haixing Dai, Hossein Mansour, Jiangfan Zhang, Shlok Kumar Mishra, Wei Sun, Xuanming Cui, Yanli Liu, Qi Guo, Max Xiangjun Fan, Jun Xiao

arXiv 2610.11316首次发表:更新:

发表机构

Meta AI(Meta AI)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出MetaEncoder,将预训练的Muse-Glimmer 30B解码器微调为遵循指令的决策双编码器,经11个基准套件、190项任务评估,其在多模态决策等部分场景优于SOTA多模态编码器,同时明确了自身在推理密集型任务上的局限。

AI 中文摘要

System One模型输出受约束的决策与概率分布,而非自由文本生成。现有主流范式依赖结构化模式对象对状态、意图和候选选项进行编码,本文重新探讨完全基于自然语言的System One接口。在该框架中,用户请求与每个候选选项均以自然语言表达,辅以多模态(图像与视频)辅助输入。我们提出MetaEncoder,它将预训练的Muse-Glimmer 30B解码器微调为遵循指令的决策编码器。为有效适配小闭集(<256)与大规模开集(数百万)候选空间,MetaEncoder采用通过单向对比学习训练的双编码器架构,用于请求-候选对齐。我们在11个基准套件、190项任务上开展广泛评估,涵盖多模态决策、闭集理解与开集检索,结果凸显MetaEncoder在部分场景优于SOTA多模态编码器,同时也明确其在推理密集型任务上的当前局限。

英文摘要

System One models output constrained decisions and probability distributions rather than free-form text generation. While prevailing paradigms rely on structured schema objects to encode state, intent, and candidate choices, we revisit a fully natural language-based System One interface. In this framework, both the user request and each candidate option are expressed in natural language, supported by multimodal (image and video) auxiliary inputs. We introduce MetaEncoder, which fine-tunes a pre-trained Muse-Glimmer 30B decoder into an instruction-following decision-making encoder. To scale effectively across both small closed-set (< 256) and massive open-set (millions) candidate spaces, MetaEncoder employs a bi-encoder architecture trained via unidirectional contrastive learning for request-candidate alignment. We conduct extensive evaluations across 11 benchmark suites and 190 tasks spanning multimodal decision-making, understanding (closed-set) and retrieval (open-set), highlighting where MetaEncoder beats SOTA multimodal encoders, as well as its current limits on reasoning-intensive tasks.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑