This dork’s ears might not work. I say SPECIFICALLY in the clip Qwen already used linear attention, I know it’s not a new idea. But they are such a fucking larping idiot that they completely glosses over the fact that Kimi replace’s gloss scalar forget gate with a channel wise one so every feature dimension gets its own forget rate and the quality improvement is obvious. As well as failing to mention it’s always on max reasoning… We aren’t even go into all the trade offs with storing a dense recurrent state and how often you need to do it to avoid prefill or how it lacking sparsity puts more pressure on BW requirements making flash less appealing. I’m sorry I can’t explain everything to you ! Maybe spend some time to try to actually build something or educate yourself so your stupid little brain can actually contribute to the conversation other than reacting like a fucking goy to every pieces of news that drops and the. deciding to waste your time commenting on my podcast clips. Fuck you and your ancestors back 18 generations. Sincerely, Bubbleboi