Evidence-backed AMD Strix Halo local-AI setup and benchmarks: Qwen3.8, Ollama, llama.cpp, Vulkan/ROCm, large GGUFs, and cross-OEM results.
Tags
Similar apps
More in this categoryODS V3 Pre-Release: Public testing and refinement ahead of the official V3 launch. Turn your PC, Mac, or Linux box into a private AI server.
7.2kPythonApache-2.0Self-hostedllama.cpp/ik_llama.cpp launcher: loads big MoE models across mismatched multi-GPU rigs by exact VRAM math.
281GoMITSelf-hostedA local LLM server built for concurrent work. Drop-in replacement for Ollama and/or OpenAI and Ollama APIs on one port.
189RustPrivacy first, AI meeting assistant with 4x faster Parakeet/Whisper live transcription, speaker diarization, and Ollama summarization built on Rust.
32kRustMITWindowsmacOSDesktopFine-tune LLMs from one YAML. Layer streaming trains an 8B model on a 4 GB laptop GPU.
8.5kPythonApache-2.0CLIGemma 4 26B-A4B inference in ~2 GB of RAM on any M-series MacBook
6.9kSwiftApache-2.0macOSDesktop