Native MLX runtime for Laya typed decision models — 7–14 ms short decisions on M3 Max. No text generation, PyTorch, or cloud API.
Tags
Similar apps
More in this categoryGemma 4 26B-A4B inference in ~2 GB of RAM on any M-series MacBook
6.9kSwiftApache-2.0macOSDesktopThe fastest way to run Qwen 3.8 Flash Next, Qwen 3.8 27B and Ternary Bonsai 2 27B on a Mac: 125 tok/s in OpenCode on an M5 Max, and a 27B model on 16 GB Macs.
2.6kPythonApache-2.0macOSDesktopThe open-source System One decision model. Sub-15ms, non-autoregressive, local drop-in alternative to TypeSafe Jev.
876PythonApache-2.0Run a 105 GB AI model on a Mac that can't hold it. Slotstream streams Qwen3.8-Flash-Next (125B mixture of experts) from your SSD and caches the busiest experts…
453SwiftMITmacOSDesktopAn easy-to-use, fast toolkit to scale up RL post-training on a single node.
323PythonApache-2.0Self-hosted