chadhurley25075-png/pd-bridge
pd-bridge enables heterogeneous inference by splitting prefill and decode stages across different hardware. It routes CUDA prefill on DGX Spark to Metal decode on Mac Studio via 10GbE. The project supports DeepSeek-V4-Flash and uses vLLM with oMLX. This setup allows developers to leverage specialized hardware for different inference phases.
README
chadhurley25075-png/pd-bridge View on GitHub
Loading the README from GitHub…