SNIP++: 用于符号回归的细粒度符号-数值对齐
SNIP++: Fine-Grained Symbolic-Numerical Alignment for Symbolic Regression
浏览论文内容
中文总结 AI 辅助
针对符号回归中多模态嵌入仅全局对齐、无法感知局部编辑影响的问题,提出SNIP++方法,通过结构位置编码与多粒度对比学习实现细粒度符号-数值对齐,显著提升局部编辑区分能力并迁移至外部语料库。
中文摘要 AI 辅助
数学表达式及其产生的数值行为是同一底层函数的两种视图,连接二者是科学发现的核心。符号回归(SR)直接依赖这种连接:它搜索能够重现给定行为的表达式。近年来的多模态模型通过在共享表示空间中对符号表达式及其数值行为进行嵌入来学习这种连接。我们表明,这种嵌入空间仅实现了全局对齐:完整表达式对应完整行为,但表达式各部分的贡献并未被表示。这种粒度差距使得模型无法判断对表达式的局部编辑如何改变其行为,而这正是SR中的核心操作。我们提出了一种弥合这一差距的组成性对齐方法:结构位置编码向编码器暴露表达式的子结构,多粒度对比目标将每个子表达式在其产生的行为中进行基础化,然后将这种基础化传播到完整表达式。由此产生的表示在很大程度上弥合了符号嵌入与数值嵌入之间的模态差距,能够可靠地区分原始对齐无法区分的局部编辑效果,并可迁移到外部SR语料库。
英文摘要
Mathematical expressions and the numerical behavior they produce are two views of the same underlying function, and connecting them is central to scientific discovery. Symbolic Regression (SR) relies on this connection directly: it searches for an expression that reproduces a given behavior. Recent multi-modal models learn this connection by embedding symbolic expressions and their numerical behavior in a shared representation space. We show that this embedding space is only globally aligned: complete expressions correspond to complete behaviors, but the contribution of individual parts of an expression is not represented. This granularity gap leaves the model unable to tell how a local edit to an expression changes its behavior, the central operation in SR. We introduce a compositional alignment method that closes this gap: a structural positional encoding exposes the substructure of an expression to the encoder, and a multi-granularity contrastive objective grounds each subexpression in the behavior it produces before propagating this grounding to the full expression. The resulting representations close much of the modality gap between symbolic and numerical embeddings, reliably distinguish the effects of local edits that the original alignment cannot, and transfer to external SR corpora.
发表机构
- Université Laval(拉瓦尔大学)
- Mila(米拉研究所)
- Carnegie Mellon University(卡内基梅隆大学)
机构由 AI 辅助整理,请以论文原文为准。