Var-JEPA:联合嵌入预测架构的变分公式——弥合预测式与生成式自监督学习
文章背景与核心概要
联合嵌入预测架构(JEPA)长期以来被视为基于似然的自监督学习的一种非生成式替代方案,其核心在于表征空间内的预测,而非观测空间内的重构。然而,本文指出预测建模与概率生成建模之间的界限更多是修辞上的而非结构上的:标准的 JEPA 设计直接映射了耦合潜变量模型在变分推断下得到的变分后验与学习条件先验,从而揭示了标准 JEPA 实质上是一种确定性的特化形式。
基于这一深刻洞察,作者提出了变分 JEPA(Var-JEPA)。该方法通过优化单一的证据下界(ELBO),显式构建了潜变量生成结构。这不仅在无需特定防坍塌正则化项的情况下学到了高质量的表征,还实现了潜变量空间中有据可依的不确定性量化。针对表格数据实例化的 Var-T-JEPA 在各项真实世界表格基准测试中,表现出了优于 T-JEPA 的表征学习与下游任务性能,同时保持了对强大原始特征基准的强劲竞争力。
Var-JEPA: A Variational Formulation of the Joint-Embedding Predictive Architecture - Bridging Predictive and Generative Self-Supervised Learning
Summary
The Joint-Embedding Predictive Architecture (JEPA) is traditionally viewed as a non-generative alternative to likelihood-based self-supervised learning, focusing on predictions within a representation space rather than reconstructions in the observation space.
This paper argues that the division between predictive and probabilistic generative modeling is primarily rhetorical rather than structural: - The standard JEPA design—featuring coupled encoders with a context-to-target predictor—directly mirrors variational posteriors and learned conditional priors derived from variational inference in coupled latent-variable models. - Standard JEPA acts as a deterministic specialization where regularization relies on architectural and training heuristics rather than an explicit likelihood.
Building on this insight, the authors introduce Variational JEPA (Var-JEPA), which makes the latent generative structure explicit by optimizing a single Evidence Lower Bound (ELBO). This formulation yields meaningful representations without ad-hoc anti-collapse regularizers while enabling principled uncertainty quantification in the latent space. Instantiated for tabular data as Var-T-JEPA, the framework achieves superior representation learning and downstream performance across real-world tabular benchmarks compared to T-JEPA, while staying competitive with strong raw-feature baselines.
文档元数据
Document Metadata
| 字段 | 详情 |
|---|---|
| arXiv ID | arXiv:2603.20111 [cs.LG] |
| 作者 | Moritz Gögl, Christopher Yau |
| 学科分类 | 机器学习 (cs.LG); 人工智能 (cs.AI) |
| 发表/会议 | 第43届国际机器学习会议论文集,韩国首尔 (PMLR 306, 2026) |
| 提交日期 | 2026年3月20日 (v1);最后修订于 2026年8月28日 (v2) |
Field Details arXiv ID arXiv:2603.20111[cs.LG]Authors Moritz Gögl, Christopher Yau Subjects Machine Learning ( cs.LG); Artificial Intelligence (cs.AI)Published / Venue Proceedings of the 43rd International Conference on Machine Learning, Seoul, South Korea (PMLR 306, 2026) Submitted Date March 20, 2026 (v1); Last revised August 28, 2026 (v2)
摘要
Abstract
联合嵌入预测架构(JEPA)通常被看作是基于似然的自监督学习的一种非生成式替代方案,它强调在表征空间中进行预测,而不是在观测空间中进行重构。我们认为,由此产生的与概率生成建模的分离在很大程度上是修辞上的,而非结构上的:典型的 JEPA 设计(具有上下文到目标预测器的耦合编码器)镜像了将变分推断应用于特定类别的耦合潜变量模型时获得的变分后验和学到的条件先验,并且标准 JEPA 可以被视为一种确定性的特化,其中正则化是通过架构和训练启发式而非显式似然施加的。基于这一观点,我们推导出了变分 JEPA(Var-JEPA),它通过优化单一的证据下界(ELBO)使潜变量生成结构显式化。这无需特殊的防坍塌正则化器即可产生有意义的表征,并允许在潜变量空间中进行有原则的不确定性量化。我们将该框架实例化用于表格数据(Var-T-JEPA),并获得了强大的表征学习和下游性能,在真实的表格基准测试中超越了 T-JEPA,同时与强大的原始特征基准保持了竞争力。
The Joint-Embedding Predictive Architecture (JEPA) is often seen as a non-generative alternative to likelihood-based self-supervised learning, emphasizing prediction in representation space rather than reconstruction in observation space. We argue that the resulting separation from probabilistic generative modeling is largely rhetorical rather than structural: the canonical JEPA design (coupled encoders with a context-to-target predictor) mirrors the variational posteriors and learned conditional priors obtained when variational inference is applied to a particular class of coupled latent-variable models, and standard JEPA can be viewed as a deterministic specialization in which regularization is imposed via architectural and training heuristics rather than an explicit likelihood. Building on this view, we derive the Variational JEPA (Var-JEPA), which makes the latent generative structure explicit by optimizing a single Evidence Lower Bound (ELBO). This yields meaningful representations without ad-hoc anti-collapse regularizers and allows principled uncertainty quantification in the latent space. We instantiate the framework for tabular data (Var-T-JEPA) and achieve strong representation learning and downstream performance, improving over T-JEPA across real-world tabular benchmarks while remaining competitive with strong raw-feature baselines.
访问链接与资源
Access Links & Resources
- PDF 文档: 查看 PDF
- HTML 版本(实验性): arXiv HTML 版本
- DOI 链接: 10.48550/arXiv.2603.20111
- 开源许可: 知识共享署名 4.0 国际许可协议

- PDF: View PDF
- HTML (Experimental): arXiv HTML version
- DOI: 10.48550/arXiv.2603.20111
- License: Creative Commons Attribution 4.0 International