Really good analysis from the team at @SemiAnalysis_ and @dylan522p love seeing the game evolve. After reading this article it begs the question; how will compute buyers react, do they: 1/ use a more flexible GPU cluster with tools like TileRT to basically enable expensive turbo charged 'tokens on demand' batch size 1 without dedicated decode silicon capex or do they: 2/ do what we are doing @fastinference and strap dedicated decode silicon next to GPUs and have a team unlocking exceptional fast token economics (although with more capex and team energy needed) I guess option 1 will work for smaller clouds to offer premium tokens with bad margins so they can stay part of the 'fast benchmarks', but Option 2 will be where the labs, or inference clouds like @FireworksAI_HQ invest. Decode silicon will be key differentiator and show up in economies of scale that only the big boys will be able to operate. We are building @fastinference to basically take on the capex outlay for decode silicon so the clouds can rent it off of us speeding up adoption as we saw last year that this was where the market would head. Thoughts?