arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.23242cs.MM

留意沙发!通过弱到强任务向量注入激发多模态大语言模型在室内设计中的推理能力

Mind the Couch! Eliciting MLLM Reasoning in Interior Design via Weak-to-Strong Task Vector Injection

Yuxuan Yang, Jingyao Wang, Luntian Mou

首次发表
浏览论文内容

中文总结 AI 辅助

针对MLLMs在室内设计中因模态对齐问题产生的空间碰撞与审美不和谐,本文提出DART-I机制,通过弱到强任务向量注入引导MLLMs推理,无需微调即可实现精确推理,且经实验验证有效。

中文摘要 AI 辅助

多模态大语言模型(MLLMs)已展现出出色性能,但在面临室内设计中受密集约束的空间时,常出现严重的模态对齐问题。由于视觉编码过程中丢失了高频局部拓扑细节和细粒度审美变化,现有MLLMs常出现幻觉,产生物理空间碰撞和视觉审美不和谐。为解决该问题,本文提出双先验激活残差任务向量注入机制(DART-I),用于MLLMs,将范式从有损文本提示转向直接潜在干预,利用弱到强确定性规则锚定MLLMs在室内设计中的因果推理。具体而言,DART-I分三步运行:首先,用极轻量的弱专家从图像中显式提取连续空间距离和色彩排版特征;随后,通过线性投影网络将这些确定性先验转换为定向任务向量;最后,将这些向量作为残差项动态注入冻结MLLMs的潜在空间,引导MLLMs进行精确的室内设计推理。该方法跳出了传统范式,无需对MLLMs进行微调即可实现精确推理,有效规避了高昂的计算成本和灾难性遗忘。在多个基准上开展的大量实验证明了DART-I的有效性和优势。

英文摘要

Multimodal Large Language Models (MLLMs) have demonstrated great performance, yet they often suffer from severe modality misalignment when confronted with densely constrained spaces for interior design. Due to the loss of high-frequency local topological details and fine-grained aesthetic shifts during visual encoding, existing MLLMs frequently hallucinate, yielding physical spatial collisions and visual aesthetic dissonance. To address this, we propose Dual-prior Activation Residual Task-vectors Injection mechanism (DART-I) for MLLMs. It shifts the paradigm from lossy text-prompting to direct latent intervention, utilizing weak-to-strong deterministic rules to anchor the causal reasoning of MLLMs for interior design. Specifically, DART-I operates in three steps: it first explicitly extracts continuous spatial distance and color typography features from images using extremely lightweight weak experts; subsequently, it transforms these deterministic priors into directional task vectors via a linear projection network; these vectors are dynamically injected as residual terms into the latent space of the frozen MLLMs, steering MLLMs towards precise reasoning for interior design. Stepping outside the conventional paradigms, our method achieves precise reasoning without fine-tuning the MLLMs, effectively bypassing expensive computational costs and catastrophic forgetting. Extensive experiments on various benchmarks demonstrate the effectiveness and advantages of DART-I.

↑