Speculative decode avoids a lot of the pain in serving linear attention mechanism especially at low batch. Surprised they were able to draft this ontop of kimi that fast too.