Run a 105 GB AI model on a Mac that can't hold it. Slotstream streams Qwen3.8-Flash-Next (125B mixture of experts) from your SSD and caches the busiest experts…
Run a 105 GB AI model on a Mac that can't hold it. Slotstream streams Qwen3.8-Flash-Next (125B mixture of experts) from your SSD and caches the busiest experts in memory, so it runs on Macs with 16 to 64 GB. One native Swift binary on MLX and Metal, no Python, offline. Works with Claude Code, Codex and Ollama or OpenAI clients.
Tags
Similar apps
More in this categoryThe fastest way to run Qwen 3.8 Flash Next, Qwen 3.8 27B and Ternary Bonsai 2 27B on a Mac: 125 tok/s in OpenCode on an M5 Max, and a 27B model on 16 GB Macs.
2.6kPythonApache-2.0macOSDesktopUp to 4× faster LLM decoding on Apple Silicon, lossless. Native MLX port of DeepSeek's DSpark & z-lab's DFlash speculative decoding — Gemma-4, Qwen3.8,…
702PythonMITmacOSDesktopGemma 4 26B-A4B inference in ~2 GB of RAM on any M-series MacBook
6.9kSwiftApache-2.0macOSDesktopHigh-performance OpenAI and Anthropic compatible LLM inference server for Apple Silicon.
1.6kPythonApache-2.0macOSDesktopYour iPhone helps your Mac run a 27B model: faster prompt reading and more context over a USB-C cable
920PythonMITmacOSDesktopiOS