Internet-scale training: 10B now, then 100B, then a trillion parameters
Greg Osuri - Akash: Decentralized Compute Marketplace - ep 158 (Zima Red)
“The first billion parameter model is being trained now on Nous — by the way, using Akash. First 10 billion parameter is being trained by Prime Intellect, all using Akash… Next model is a 100 billion parameter model, and after that you’re going to see a trillion parameter model… With enough capital and with enough incentives, I don’t see a reason why training over the internet can be a profitable venture.” — 00:37:27
Context: Nous Research and Prime Intellect had just cracked heterogeneous, non-colocated training (DiLoCo/OpenDiLoCo); he notes GPT-4 is believed to be ~1.2T parameters as the target scale.