Serving LLMs on Tenstorrent Hardware: Inside the vLLM TT Plugin
vLLM introduces a plugin enabling Tenstorrent accelerator support via its out-of-tree mechanism. The integration preserves the standard OpenAI-compatible API and request formats for client applications. It supports Llama, Qwen, Mistral, and Gemma architectures by mapping them to TT-Metal runtime implementations. Multimodal models are included in this initial release scope.