Rapid-MLX is an open-source (Apache 2.0) OpenAI- and Anthropic-compatible LLM inference server for Apple Silicon, built on MLX, focused on reliable tool…
Rapid-MLX is an open-source (Apache 2.0) OpenAI- and Anthropic-compatible LLM inference server for Apple Silicon, built on MLX, focused on reliable tool calling for coding agents. Up to 4× faster than Apple's MLX (mlx-lm) on the same weights.
Open-source alternative to
Tags
Similar apps
More in this categoryHigh-performance OpenAI and Anthropic compatible LLM inference server for Apple Silicon.
1.6kPythonApache-2.0macOSDesktopThe fastest way to run Qwen 3.8 Flash Next, Qwen 3.8 27B and Ternary Bonsai 2 27B on a Mac: 125 tok/s in OpenCode on an M5 Max, and a 27B model on 16 GB Macs.
2.6kPythonApache-2.0macOSDesktopNative LLM inference server for Apple Silicon. OpenAI + Anthropic API compatible. No Python.
1.8kZigmacOSDesktopA local inference engine for Apple silicon, built around the model.
1.3kC++Apache-2.0macOSDesktopUp to 4× faster LLM decoding on Apple Silicon, lossless. Native MLX port of DeepSeek's DSpark & z-lab's DFlash speculative decoding — Gemma-4, Qwen3.8,…
702PythonMITmacOSDesktopRun a 105 GB AI model on a Mac that can't hold it. Slotstream streams Qwen3.8-Flash-Next (125B mixture of experts) from your SSD and caches the busiest experts…
453SwiftMITmacOSDesktop