arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.25138cs.LGcs.AI

漂移变分自编码器:通过条件后验流匹配统一生成与表示学习

Drift Variation Autoencoder: Unifying Generation and Representation Learning through Conditional Posterior Flow Matching

Jiarui Cao

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出漂移变分自编码器,通过条件后验流匹配统一生成与表示学习,在CrossGeom-4基准上实现了高表示精度与模态一致性,完成多模态概念验证。

中文摘要 AI 辅助

随机掩码、裁剪或模态去除使得确定性重构成为不完整目标:一个观测值可对应多个干净补全结果。本研究将对应后验分布 $P(X\mid C)$ 作为条件生成与生成式充分表示学习的共同统计对象。漂移变分自编码器训练掩码编码器 $Z=E(C)$ 和条件流解码器,采用单干净预测流匹配损失。分析首先将理想条件KL散度分解为生成器近似项和表示缺陷项 $I(X;C\mid Z)$,随后推导条件流匹配的正交风险分解。对于仿射高斯路径,当且仅当 $P(X\mid Z)=P(X\mid C)$ 时,干净预测表示间隙为零。因此,流匹配诱导的依赖编码器的超额干净预测风险与轮廓化理想条件KL散度具有相同的后验充分零集,但并非数值相等的目标。零噪声端点的精确条件场在联合理想最优时生成 $P(X\mid Z)$,进而生成 $P(X\mid C)$。当完整模态元组保持为每个观测掩码的流目标时,该结果可扩展至连续多模态乘积空间。在18次运行的受控基准CrossGeom-4上,可观测因子的线性探针 $R^2$ 为0.9990-0.9992;打乱联合模型编码器的条件使条件误差增加13.5倍-15.7倍;与独立目标解码器相比,联合目标注意力使两个输出共享的未观测因子的不一致性降低90.1%-92.8%。可见模态也被生成和重构,直接验证了完整元组目标。无条件模式平衡仍不完善,因此经验性结论仅限于受控多模态概念验证。

英文摘要

Stochastic masking, cropping, or modality removal makes deterministic reconstruction an incomplete target: one observation can admit many clean completions. This work takes the corresponding posterior $P(X\mid C)$ as the common statistical object for conditional generation and generatively sufficient representation learning. Drift Variation autoencoder trains a masked encoder $Z=E(C)$ and a conditional flow decoder with one clean-prediction Flow Matching loss. The analysis first decomposes the ideal conditional KL into generator approximation and the representation deficiency $I(X;C\mid Z)$. It then derives orthogonal risk decompositions for conditional Flow Matching. For an affine Gaussian path, the clean-prediction representation gap is zero if and only if $P(X\mid Z)=P(X\mid C)$. Thus the encoder-dependent excess clean-prediction risk induced by Flow Matching and the profiled ideal conditional KL have the same posterior-sufficient zero set, without being numerically equal objectives. An exact conditional field with a zero-noise endpoint then generates $P(X\mid Z)$ and hence $P(X\mid C)$ at a joint ideal optimum. The result extends to continuous multimodal product spaces when the complete modality tuple remains the Flow target for every observation mask. On CrossGeom-4, an 18-run controlled benchmark, observable factors have linear-probe $R^2$ of $0.9990$-$0.9992$, shuffling the joint model's encoder condition increases conditional error by $13.5\times$-$15.7\times$, and joint target attention reduces disagreement on an unobserved factor shared by two outputs by $90.1$-$92.8\%$ relative to independent target decoders. Visible modalities are also generated and reconstructed, directly validating the full-tuple objective. Unconditional mode balance remains imperfect, delimiting the empirical claim to a controlled multimodal proof of concept.

发表机构

  • The Chinese University of Hong Kong(香港中文大学)

机构由 AI 辅助整理,请以论文原文为准。

↑