arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过上下文工程进行隐式微调:多模态实体对齐的课程学习框架

Implicit Fine-tuning via Context Engineering: A Curriculum Learning Framework for Multimodal Entity Alignment

Yunpeng Hong, Chenyang Bu, Di Wu, Yi He, Xindong Wu

arXiv 2607.10532首次发表:更新:

AI 中文总结

研究多模态实体对齐问题,提出PTFEA课程学习框架,通过自适应难度调制和三阶段渐进推理将微调策略转化为上下文工程,实验证明该框架性能优越,提供了统一上下文工程和微调的理论框架。

AI 中文摘要

多模态实体对齐旨在识别不同模态中的等效实体。现有方法通过黑箱上下文工程策略提高性能,但依赖大语言模型参数能力且缺乏理论可解释性。为此,首先从理论上验证了上下文工程与模型微调在多模态实体对齐任务中的数学等价性。在此基础上,提出了PTFEA框架,它将微调策略转化为可解释的上下文工程。具体包括自适应难度调制和三阶段渐进推理。在五个公共数据集上的实验表明,PTFEA始终优于强基线,还大幅降低了运行时间和令牌消耗。这项工作提供了首个统一上下文工程和微调的理论框架,为未来研究奠定基础。

英文摘要

Multimodal Entity Alignment (MMEA) aims to identify equivalent entities across different modalities. While existing methods enhance MMEA performance through black-box context engineering strategies, their reliance on LLM parameter capacity and lack of theoretical interpretability remain unresolved. To this end, we first theoretically validate the mathematical equivalence between context engineering and model fine-tuning in MMEA tasks, demonstrating that prompt components simulate contrastive learning-based sequential fine-tuning in MMEA. Building on this foundation, we then propose PTFEA, a curriculum-learning-inspired framework that translates fine-tuning strategies into interpretable context engineering. Specifically, adaptive difficulty modulation dynamically adjusts information injection stages using confidence thresholds, establishing mathematical equivalence between curriculum learning weights and context sample selection; and three-stage progressive inference incorporates entity information from simple to complex cases, mirroring the gradient descent process in fine-tuning. Experiments on five public datasets demonstrate that PTFEA consistently outperforms strong baselines. In particular, on the ICWIKI dataset, PTFEA narrows the H@1 gap between Qwen2.5-72B and 14B to 0.6%. Moreover, compared with the representative context-engineering-based MMEA method MM-ChatAlign, PTFEA reduces the runtime of Qwen2.5-72B from 21 hours to 1 hour and lowers token consumption from 2200-3000 to 200-400, achieving over 80% reduction on the ICWIKI dataset. This work provides the first theoretical framework unifying context engineering and fine-tuning in MMEA, paving the way for future research that seeks to translate additional fine-tuning strategies into context engineering paradigms. Our code is available at https://github.com/DMiC-Lab-HFUT/PTFEA.

CommentsAccepted by KDD 2026

Journal refSIGKDD 2026

DOI:10.1145/3770855.3817732

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑