A local LLM server built for concurrent work. Drop-in replacement for Ollama and/or OpenAI and Ollama APIs on one port.
A local LLM server built for concurrent work. Drop-in replacement for Ollama and/or OpenAI and Ollama APIs on one port. Requests that share a prompt reuse each other's KV cache instead of each prefilling it. Rust, wrapping llama.cpp.
Open-source alternative to
Tags
Similar apps
More in this categoryllama.cpp/ik_llama.cpp launcher: loads big MoE models across mismatched multi-GPU rigs by exact VRAM math.
281GoMITSelf-hostedLocal Qwen 3.8 27B uncensored Q4_K_M + harvested SYSTEM pack. Official 3.8 weights, not a 3.6 retitle.
337JavaScriptA Web Interface for chatting with your local LLMs via the ollama API
1.3kTypeScriptMITSelf-hostedThe fastest way to run Qwen3.8-Flash-Next on Strix Halo (gfx1151)
915ShellRun a 105 GB AI model on a Mac that can't hold it. Slotstream streams Qwen3.8-Flash-Next (125B mixture of experts) from your SSD and caches the busiest experts…
453SwiftMITmacOSDesktop