OpenAppsSubmit

Halogen Flash Server

by peonist-ai

The fastest way to run Qwen3.8-Flash-Next on Strix Halo (gfx1151)

Halogen Flash Server on GitHub

Tags

  • Run a 105 GB AI model on a Mac that can't hold it. Slotstream streams Qwen3.8-Flash-Next (125B mixture of experts) from your SSD and caches the busiest experts…

    453SwiftMITmacOSDesktop
    1mo ago
  • FoxAI & LLMalt. to Ollama

    A local LLM server built for concurrent work. Drop-in replacement for Ollama and/or OpenAI and Ollama APIs on one port.

    189Rust
    7mo ago
  • ODS V3 Pre-Release: Public testing and refinement ahead of the official V3 launch. Turn your PC, Mac, or Linux box into a private AI server.

    7.2kPythonApache-2.0Self-hosted
    8mo ago
  • Local-first healthcare AI: clinical NER and HIPAA PII de-identification on hardware you control.

    5.5kPythonApache-2.0macOSDesktopAndroid
    1y ago
  • Personal AI Notebooks. Organize files & webpages and generate notes from them. Open source, local & open data, open model choice (incl. local).

    3.6kTypeScriptApache-2.0
    12mo ago
  • The fastest way to run Qwen 3.8 Flash Next, Qwen 3.8 27B and Ternary Bonsai 2 27B on a Mac: 125 tok/s in OpenCode on an M5 Max, and a 27B model on 16 GB Macs.

    2.6kPythonApache-2.0macOSDesktop
    5mo ago