
3 min read
Zyphra shows LLM plasticity loss scales sublinearly,…
Zyphra Research found GPT-style transformers from 5M to 314M params lose plasticity during continual and even stationary training. Bigger models delay the cliff but scaling law gains are sublinear.