arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

柏拉图式任务算术

Platonic Task Arithmetic

Junghwan Park, Woojin Cho

arXiv 2610.00929首次发表:更新:

发表机构

TelePIX(TelePIX)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出柏拉图式任务向量与通用任务描述符,通过最小二乘或低秩适配器实现跨异构模型的任务知识迁移与组合,保留74-80%性能。

AI 中文摘要

针对同一任务特化的模型会收敛到相似的行为,但产生这种行为的参数更新却缺乏共同的坐标系,因此权重空间中的任务算术只能局限于单个模型,无法在没有结构对应关系的情况下跨架构进行。借鉴柏拉图的洞穴寓言,我们假设这些模型特定的更新是一个共享的、与模型无关的对象的影子,我们称之为柏拉图式任务向量。为了使其适用于将图像或音频编码器与文本编码器配对的模型,我们引入了通用任务描述符:其形状独立于架构和嵌入维度的矩阵,这些矩阵记录任务的功能效果,并支持作为矩阵运算的加法和取反。将描述符转移到目标模型中意味着编辑目标模型,直到它在任务的无标签探针图像和类名提示上重现该描述符,这不需要逐图像标签。我们通过两种方式实现这种编辑。首先,描述符分解为图像嵌入上的位移场,因此一次最小二乘求解即可得到一个线性算子,该算子作为权重编辑折叠到目标的最后一层;通过线性性质,一组这样的算子允许以任何强度进行任何组合作为带符号和。其次,在相同目标上训练的低秩适配器可以到达每一层并联合拟合组合,代价是每次编辑进行一次优化。异构模型仅部分共享此对象,模型特定的残差在范数上可与共享组件相当,但跨模型迁移仍然保留了目标自身描述符收益的74-80%。在六个模型族、八个分类任务和一个音频-文本设置中的实验表明,在两种实现下,任务知识都能跨异构模型迁移和组合。

英文摘要

Distinct pre-trained models specialized for the same task converge to closely similar behavior, yet the parameter updates that produce it share no common coordinate system. Weight-space task arithmetic is therefore confined to a single model, and transporting an update between models requires a structural correspondence. Drawing on Plato's allegory of the cave, we hypothesize that these model-specific updates are shadows cast by one shared, model-agnostic object, the platonic task vector. To make it operational across models of different architectures, we introduce Universal Task Descriptors, matrices whose shape is independent of architecture and embedding dimension, which record a task's functional effect and admit addition and negation as ordinary matrix operations, and we transfer a descriptor into a target in two ways. A single least-squares solve returns a linear operator folded into the target's last layer, and a bank of such operators, one per source and task, realizes any composition as a signed sum of its entries. Alternatively, a low-rank adapter of the target's encoder is trained on the same objective at the price of one optimization per edit. Despite a model-specific residual comparable in norm to the shared component, transfer from another model retains 74 to 80 percent of the gain the target's own descriptors attain. Experiments across six model families, eight tasks and audio-text models confirm both realizations.

CommentsNeurIPS2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑