OpenAppsSubmit

Vllm Mlx

by waybarrios

AI & LLMmacOSDesktop

High-performance OpenAI and Anthropic compatible LLM inference server for Apple Silicon.

High-performance OpenAI and Anthropic compatible LLM inference server for Apple Silicon. Native MLX, continuous batching, multimodal models, MCP tool calling, and Claude Code support.

Vllm Mlx on GitHub

Tags

  • Rapid MLXAI & LLMalt. to Lm Studio

    Rapid-MLX is an open-source (Apache 2.0) OpenAI- and Anthropic-compatible LLM inference server for Apple Silicon, built on MLX, focused on reliable tool…

    4kPythonmacOSDesktop
    8mo ago
  • Run a 105 GB AI model on a Mac that can't hold it. Slotstream streams Qwen3.8-Flash-Next (125B mixture of experts) from your SSD and caches the busiest experts…

    453SwiftMITmacOSDesktop
    1mo ago
  • The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor…

    276kJavaScriptMIT
    9mo ago
  • Run Claude Code 100% on-device with local AI on Apple Silicon. MLX-native Anthropic-API server. 6 fighters incl.

    3.3kPythonMITmacOSDesktop
    7mo ago
  • The fastest way to run Qwen 3.8 Flash Next, Qwen 3.8 27B and Ternary Bonsai 2 27B on a Mac: 125 tok/s in OpenCode on an M5 Max, and a 27B model on 16 GB Macs.

    2.6kPythonApache-2.0macOSDesktop
    5mo ago
  • Native LLM inference server for Apple Silicon. OpenAI + Anthropic API compatible. No Python.

    1.8kZigmacOSDesktop
    8mo ago