25 years
on one chip.Nine days on a datacentre.
Hardware buys back the time, but the money and the carbon get spent either way. This is what one training run really costs.
1.02K H100 SXMs finish GPT-3 (175B) in 9 days.
On a single chip, the same run would take 25.2 years.
However many chips share the work, the bill stays the same. It just arrives sooner.
Put it into perspective
48 return flights London to New York
10 years of one UK person's territorial emissions
88 UK homes powered for a year
Estimates assume 40% real-world chip utilisation held whatever the chip count, so wall-clock time scales perfectly with the number of chips, which no real cluster manages. Electricity is the chip draw multiplied by 1.4 for the rest of the node, then by a further 1.1 for the building. The figures are illustrative and approximate. Provided by the Institute of AI for interest and learning.
Four ideas behind
the floor.
The 6ND rule
Training a dense transformer costs roughly six floating-point operations per parameter, per token. That single formula sets the scale of everything on this page.
Data matters as much as size
The Chinchilla result showed that many models were undertrained. For a fixed compute budget, balancing parameters and tokens beats simply making the model bigger.
More chips, same bill
Adding GPUs makes training faster but not cheaper: the total compute is fixed, so the energy, carbon and rental cost stay roughly the same however many chips share the work.
Where the power comes from
The same run can emit many times more carbon on a fossil-heavy grid than on a low-carbon one. Where and when you train is a design decision, not an afterthought.
Membership is free. Accreditation is the standard.
Keep
exploring.
Get AI is a collection of games, experiences, and tools that make AI easier to understand from the Institute of AI.

