OpenAppsSubmit

Mlx Dspark

by ARahim3

AI & LLMmacOSDesktop

Up to 4× faster LLM decoding on Apple Silicon, lossless. Native MLX port of DeepSeek's DSpark & z-lab's DFlash speculative decoding — Gemma-4, Qwen3.8,…

Up to 4× faster LLM decoding on Apple Silicon, lossless. Native MLX port of DeepSeek's DSpark & z-lab's DFlash speculative decoding — Gemma-4, Qwen3.8, Muse-Glimmer, Nemotron, LFM2.5, Ornith-1.0, ternary Bonsai-27B.

Mlx Dspark on GitHub

Tags

  • The fastest way to run Qwen 3.8 Flash Next, Qwen 3.8 27B and Ternary Bonsai 2 27B on a Mac: 125 tok/s in OpenCode on an M5 Max, and a 27B model on 16 GB Macs.

    2.6kPythonApache-2.0macOSDesktop
    5mo ago
  • Run a 105 GB AI model on a Mac that can't hold it. Slotstream streams Qwen3.8-Flash-Next (125B mixture of experts) from your SSD and caches the busiest experts…

    453SwiftMITmacOSDesktop
    1mo ago
  • Rapid MLXAI & LLMalt. to Lm Studio

    Rapid-MLX is an open-source (Apache 2.0) OpenAI- and Anthropic-compatible LLM inference server for Apple Silicon, built on MLX, focused on reliable tool…

    4kPythonmacOSDesktop
    8mo ago
  • High-performance OpenAI and Anthropic compatible LLM inference server for Apple Silicon.

    1.6kPythonApache-2.0macOSDesktop
    10mo ago
  • Your iPhone helps your Mac run a 27B model: faster prompt reading and more context over a USB-C cable

    920PythonMITmacOSDesktopiOS
    8d ago
  • Persistent session memory for AI coding agents — local-first, with on-device inference, associative recall, and drift detection.

    158TypeScriptApache-2.0
    8mo ago