
The post says TPUs and other systolic-array architectures may run hybrid-attention models more efficiently and at lower cost than GPUs. It predicts the next Gemini model may use hybrid attention, potentially improving cost and intelligence per token.