Training builds the model once; inference happens every time a user makes a request, so its compute cost scales with usage. As AI products grow, inference, not training, becomes the dominant long-run demand for chips and power.