articleX25 May 2026
Every llama-server Flag Explained: The Tuning Guide For Local LLMs
witcheer@witcheer
x.com
Aithy is a private local AI runtime for state-of-the-art agent research, local inference, sandboxed tools, durable memory, and LAN Mesh resource sharing.
Run GLM-5.2 (744B MoE) on a 25GB-RAM consumer machine — pure C, zero deps, experts streamed from disk. Tiny engine, immense model. 🐦
DeepSeek 4 Flash and PRO local inference engine for Metal, CUDA and ROCm
An occasional email when notable AI dev tools and models land in the directory. No spam, unsubscribe anytime.