Yo wtf are we talking about. Thanks for again stating the obvious you absolute barnacle. Yeah the KV grows with size welcome to inference 101 you know what else makes it grow with size? The batches. You know what’s allowing you to run more batches now? Getting rid of 75% of the layers with KV. You know what else grows with batches? The state matrices? You know what 75% of layers are ? State matrices. You know what’s on the hot path? The state matrices. You know what inference involves ? Reading the state matrices 3 times for every layer and updating it. You know what’s dense unlike KV? The state matrices. You know what comes before the KV fetch? The state matrices. You know what you need for grabbing the state matrices? High bandwidth? Again what scales with batches? State matrices. Please stop stating the fucking obvious and add some nuance to the access pattern, shape, and density of what’s dominating the layers? The state matrices. State matrices say it with me now until you know how that works I don’t want to talk to you. Not sure what they taught you at Lee Kwan Yoo Brainwash University of Singapore but at my SUNY school they taught us to use our brains when a new variable is added to the equation. Please stop going back and forth with me. I have nothing but time, money, and connections. I don’t want to start a substack.