How does an AI learn how the world works — without labels, and without predicting every pixel?
In this video, we will trace 30+ years of self-supervised learning, from Becker & Hinton’s 1992 random-dot stereograms to JEPA and modern world models. We build up the key ideas one problem at a time: representation collapse, contrastive learning (CPC, MoCo, SimCLR), distillation (BYOL, DINO), masked autoencoders, and finally Joint-Embedding Predictive Architectures (I-JEPA, V-JEPA, V-JEPA 2) — models that predict in latent space and can even plan actions.
0:00 Random-dot stereogram
0:30 Learning without labels
4:19 Contrastive Learning: Positives, Negatives, and InfoNCE
8:13 Scaling Contrastive Learning: Memory Banks, MoCo, and SimCLR
14:38 Distillation: Do We Even Need Negatives?
20:12 Masked Autoencoders: Predicting Missing Pixels
23:16 I-JEPA: Predicting in Embedding Space
26:00 V-JEPA: Latent Prediction in Video
28:33 V-JEPA 2-AC: From Prediction to Planning
31:51 Preventing Collapse with Regularizing
36:27 The Endgame: LeWorldModel
source




Leave a Reply