OpenAppsSubmit

A local LLM server built for concurrent work. Drop-in replacement for Ollama and/or OpenAI and Ollama APIs on one port.

A local LLM server built for concurrent work. Drop-in replacement for Ollama and/or OpenAI and Ollama APIs on one port. Requests that share a prompt reuse each other's KV cache instead of each prefilling it. Rust, wrapping llama.cpp.

Fox on GitHub

Open-source alternative to

Tags

  • GgrunAI & LLMalt. to Ollama

    llama.cpp/ik_llama.cpp launcher: loads big MoE models across mismatched multi-GPU rigs by exact VRAM math.

    281GoMITSelf-hosted
    7mo ago
  • Local Qwen 3.8 27B uncensored Q4_K_M + harvested SYSTEM pack. Official 3.8 weights, not a 3.6 retitle.

    337JavaScript
    2mo ago
  • Go manage your Ollama models

    1.8kGoMITmacOSLinuxDesktop
    2.4y ago
  • A Web Interface for chatting with your local LLMs via the ollama API

    1.3kTypeScriptMITSelf-hosted
    3y ago
  • The fastest way to run Qwen3.8-Flash-Next on Strix Halo (gfx1151)

    915Shell
    2mo ago
  • Run a 105 GB AI model on a Mac that can't hold it. Slotstream streams Qwen3.8-Flash-Next (125B mixture of experts) from your SSD and caches the busiest experts…

    453SwiftMITmacOSDesktop
    1mo ago