OpenAppsSubmit

Slotstream

by carloslfu

AI & LLMmacOSDesktop

Run a 105 GB AI model on a Mac that can't hold it. Slotstream streams Qwen3.8-Flash-Next (125B mixture of experts) from your SSD and caches the busiest experts…

Run a 105 GB AI model on a Mac that can't hold it. Slotstream streams Qwen3.8-Flash-Next (125B mixture of experts) from your SSD and caches the busiest experts in memory, so it runs on Macs with 16 to 64 GB. One native Swift binary on MLX and Metal, no Python, offline. Works with Claude Code, Codex and Ollama or OpenAI clients.

Slotstream on GitHub

Tags

  • The fastest way to run Qwen 3.8 Flash Next, Qwen 3.8 27B and Ternary Bonsai 2 27B on a Mac: 125 tok/s in OpenCode on an M5 Max, and a 27B model on 16 GB Macs.

    2.6kPythonApache-2.0macOSDesktop
    5mo ago
  • Up to 4× faster LLM decoding on Apple Silicon, lossless. Native MLX port of DeepSeek's DSpark & z-lab's DFlash speculative decoding — Gemma-4, Qwen3.8,…

    702PythonMITmacOSDesktop
    3mo ago
  • Gemma 4 26B-A4B inference in ~2 GB of RAM on any M-series MacBook

    6.9kSwiftApache-2.0macOSDesktop
    3mo ago
  • High-performance OpenAI and Anthropic compatible LLM inference server for Apple Silicon.

    1.6kPythonApache-2.0macOSDesktop
    10mo ago
  • Your iPhone helps your Mac run a 27B model: faster prompt reading and more context over a USB-C cable

    920PythonMITmacOSDesktopiOS
    8d ago
  • Local LLM Testing & Benchmarking for Apple Silicon

    210SwiftGPL-3.0macOSDesktop
    8mo ago