the industry is finally marketing compute in tokens per second per megawatt a few years back, compute was sold in GPU-hours, you rented hardware for a set amount of compute time now the useful output is inference, and increasingly the question is how much inference you can produce from a fixed amount of power if you have 100MW, you don't really care how many GPUs sit behind the API you care how many useful tokens that 100MW can produce, at what latency, and at what cost that is the main shift I'm seeing compute-hours measure the input, tokens measure the output, tokens per second per megawatt measures how efficiently you turn the industry's scarcest resource into that output as inference becomes the dominant workload, that's a much more useful unit of account, and doesn't just apply to silicon