incoai/splash

A new local inference engine for Apple silicon is now available, written in Python and leveraging Metal. The repository supports speculative decoding and is designed for coding agents. It has gained 161 stars and 10 forks since its initial release. The project provides a direct path for running LLMs locally on macOS hardware without external cloud dependencies.

README

incoai/splash View on GitHub

Loading the README from GitHub…