Cohere's North Mini Code Megakernel Serving Engine | Cohere
Cohere publishes a serving engine for North Mini Code using a decode megakernel architecture. The system achieves 1.25 to 1.41 times faster end-to-end throughput than vLLM on a single H100 GPU. It supports continuous batching, paged attention, and ragged sequence lengths via an OpenAI-compatible endpoint. Source code is available on GitHub for developers to integrate into their inference pipelines.
