
4 min read
Self-predicting latents need exponentially less training data…
A new theory proves latent prediction recovers hierarchical structure with constant samples while token-level SSL needs exponential data. Here is what that means for JEPA, data2vec, and your pretraining budget.