A 2.78-trillion-parameter Kimi K3 running inference on a single CPU in 8.24 GB of RAM. Portable C99: no BLAS, no framework, no GPU.
Tags
Similar apps
More in this categoryEasiest and laziest way for building multi-agent LLMs applications.
3.9kPythonApache-2.0The fastest way to run Qwen 3.8 Flash Next, Qwen 3.8 27B and Ternary Bonsai 2 27B on a Mac: 125 tok/s in OpenCode on an M5 Max, and a 27B model on 16 GB Macs.
2.6kPythonApache-2.0macOSDesktopFine-tune LLMs on your Mac with Apple Silicon. SFT, DPO, GRPO, Vision, TTS, STT, Embedding, and OCR fine-tuning — natively on MLX. Unsloth-compatible API.
1.4kPythonApache-2.0macOSDesktopOffline inference engine for art, real-time voice conversations, LLM powered chatbots and automated workflows
1.3kPythonGPL-3.0DesktopSelf-hostedjevos is an open-source alternative to Jev for yes/no decisions that runs on your laptop.
1.3kC++MITSelf-hostedThe fastest way to run Qwen3.8-Flash-Next on Strix Halo (gfx1151)
915Shell