OpenAppsSubmit

LLMKube

by defilantech

AI & LLMSelf-hosted

Kubernetes operator for self-hosted LLM inference across a heterogeneous GPU fleet: NVIDIA CUDA, AMD Vulkan, and Apple Silicon Metal.

Kubernetes operator for self-hosted LLM inference across a heterogeneous GPU fleet: NVIDIA CUDA, AMD Vulkan, and Apple Silicon Metal. Runtimes: llama.cpp, vLLM, TGI, mlx-server. Multi-GPU sharding, model caching, OpenAI-compatible endpoints. Apache-2.0, run across homelab and on-prem fleets, actively developed.

LLMKube on GitHub

Tags

  • Run any local LLM engine, auto-tuned to your GPU — polished web UI + OpenAI/Anthropic-compatible API. Point Claude Code at your own machine in one command.

    281TypeScriptSelf-hostedWeb
    4mo ago
  • Your AI intranet: network the computers you already own for inference and training.

    262PythonMITSelf-hosted
    4mo ago
  • Local LLM Testing & Benchmarking for Apple Silicon

    210SwiftGPL-3.0macOSDesktop
    8mo ago
  • Native MLX runtime for Laya typed decision models — 7–14 ms short decisions on M3 Max. No text generation, PyTorch, or cloud API.

    6.9kPythonApache-2.0
    21d ago
  • Go manage your Ollama models

    1.8kGoMITmacOSLinuxDesktop
    2.4y ago
  • Native LLM inference server for Apple Silicon. OpenAI + Anthropic API compatible. No Python.

    1.8kZigmacOSDesktop
    8mo ago