← GPU Economics

Serving Apple's users would take a billion H100s — only ~700,000 exist

"Will demand for advanced AI chips (GPUs) be 1:1 for every person? We sit down with Greg Osuri to find out." (Block Fuel)

“For something like Llama 3.1 [40]5B… it takes about 24 H100s to do inference… that can serve about 20 to 30 concurrent connections at high precision… to serve Apple’s users at decent performance, you need a billion H100s. And there are about, I believe, about 700,000 H100s right now.” — 00:05:15

Context: Back-of-envelope inference math (spans into the 00:05:58 block) prompted by Apple Intelligence onboarding “a billion new devices that’ll have AI inference.”

See it among all GPU Economics predictions →