High-performance OpenAI and Anthropic compatible LLM inference server for Apple Silicon.
High-performance OpenAI and Anthropic compatible LLM inference server for Apple Silicon. Native MLX, continuous batching, multimodal models, MCP tool calling, and Claude Code support.
Tags
Similar apps
More in this categoryRapid-MLX is an open-source (Apache 2.0) OpenAI- and Anthropic-compatible LLM inference server for Apple Silicon, built on MLX, focused on reliable tool…
4kPythonmacOSDesktopRun a 105 GB AI model on a Mac that can't hold it. Slotstream streams Qwen3.8-Flash-Next (125B mixture of experts) from your SSD and caches the busiest experts…
453SwiftMITmacOSDesktopThe agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor…
276kJavaScriptMITRun Claude Code 100% on-device with local AI on Apple Silicon. MLX-native Anthropic-API server. 6 fighters incl.
3.3kPythonMITmacOSDesktopThe fastest way to run Qwen 3.8 Flash Next, Qwen 3.8 27B and Ternary Bonsai 2 27B on a Mac: 125 tok/s in OpenCode on an M5 Max, and a 27B model on 16 GB Macs.
2.6kPythonApache-2.0macOSDesktopNative LLM inference server for Apple Silicon. OpenAI + Anthropic API compatible. No Python.
1.8kZigmacOSDesktop