OpenAppsSubmit

A 2.78-trillion-parameter Kimi K3 running inference on a single CPU in 8.24 GB of RAM. Portable C99: no BLAS, no framework, no GPU.

Kimi K3 IN C on GitHub

Tags

  • Easiest and laziest way for building multi-agent LLMs applications.

    3.9kPythonApache-2.0
    2.4y ago
  • The fastest way to run Qwen 3.8 Flash Next, Qwen 3.8 27B and Ternary Bonsai 2 27B on a Mac: 125 tok/s in OpenCode on an M5 Max, and a 27B model on 16 GB Macs.

    2.6kPythonApache-2.0macOSDesktop
    5mo ago
  • Fine-tune LLMs on your Mac with Apple Silicon. SFT, DPO, GRPO, Vision, TTS, STT, Embedding, and OCR fine-tuning — natively on MLX. Unsloth-compatible API.

    1.4kPythonApache-2.0macOSDesktop
    9mo ago
  • Offline inference engine for art, real-time voice conversations, LLM powered chatbots and automated workflows

    1.3kPythonGPL-3.0DesktopSelf-hosted
    3.6y ago
  • JevAI & LLMalt. to Jev

    jevos is an open-source alternative to Jev for yes/no decisions that runs on your laptop.

    1.3kC++MITSelf-hosted
    2.2y ago
  • The fastest way to run Qwen3.8-Flash-Next on Strix Halo (gfx1151)

    915Shell
    2mo ago