arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MemeMind:用于离线上下文优化的参考引导轨迹构建

MemeMind: Reference-Guided Trace Construction for Offline Context Optimization

Run Yang, Weihang Wang, Boheng Sheng, Yuchen He, Jielei Zhang, Pengyu Chen, Zhiyu Wu, Qiang Sun, Huyang Sun, Longwen Gao

arXiv 2608.09316首次发表:更新:

发表机构

Bilibili; Fudan University(哔哩哔哩; 复旦大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

MemeMind针对离线上下文优化中失败rollout缺失成功示例的问题,通过参考引导构建工具轨迹,在梗解读任务上提升了Qwen3-VL模型的性能。

AI 中文摘要

离线上下文优化通过在保持模型冻结的同时修改智能体的指令和示例来提升智能体性能。该方法从适配集上的rollout(回滚)中学习,但部分查询仅产生失败的rollout,此时优化器无法获取可用工具得出正确答案的成功示例。我们提出MemeMind,它利用离线参考答案恢复缺失的经验:TraceBuilder识别参考所需的证据,执行文本搜索、图像检索和视觉定位,并在将生成的工具轨迹添加到适配缓冲区前进行验证;ToolGuide随后将收集的轨迹汇总为共享指南及各工具的单独指令。参考答案和构建的轨迹仅在适配阶段使用,推理阶段则使用学习到的指南及冻结的模型。我们通过动漫、漫画和游戏梗的解读研究该问题,这类梗包含编辑过的模糊视觉内容、叠加文本、长尾系列知识及特定文化参考,其解读需协调视觉定位、图像检索和文本搜索,是原生rollout组可能共同失败的高要求场景。我们在MemeX基准(含1000个由专家标注的此类梗)上评估MemeMind,在Qwen3-VL的2个模型、2个语言分区和2名独立评审下,MemeMind在GPT-5评审中,较最强的上下文优化基线,在Qwen3-VL-30B-A3B上分别提升22.0%和21.1%,在Qwen3-VL-235B-A22B上分别提升8.1%和8.0%。消融实验和保留轨迹显示,为失败组构建成功的工具使用是最大的组件增益来源,且能在推理时产生更有效的证据获取。

英文摘要

Offline context optimization improves an agent by revising its instructions and examples while keeping the model frozen. This approach learns from rollouts on an adaptation set, but some queries produce only failed rollouts. In these cases, the optimizer sees no successful example of how the available tools can reach the correct answer. We introduce MemeMind, which uses an offline reference answer to recover this missing experience. TraceBuilder identifies the evidence required by the reference, executes text search, image retrieval, and visual grounding, and verifies the resulting tool trace before adding it to the adaptation buffer. ToolGuide then summarizes the collected traces into a shared guide and separate instructions for each tool. The reference answers and constructed traces are used only during adaptation, while inference uses the learned guides with a frozen model. We study this problem through Anime, Comic, and Game meme interpretation. These memes combine edited and ambiguous visual content, overlaid text, long tail franchise knowledge, and culture specific references. Their interpretation can require coordinated visual grounding, image retrieval, and text search, making them a demanding setting in which native rollout groups may fail together. We evaluate MemeMind on MemeX, a benchmark of 1,000 such memes annotated by experts. Across two Qwen3-VL models, two language partitions, and two independent judges, MemeMind improves over the strongest context optimization baseline by 22.0% and 21.1% on Qwen3-VL-30B-A3B, and by 8.1% and 8.0% on Qwen3-VL-235B-A22B under GPT-5 judging. Ablations and held out traces show that constructing successful tool use for failed groups provides the largest component gain and produces more effective evidence acquisition at inference time.

Comments24 pages, 16 figures, 8 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑