A
Aithy
Aithy is a private local AI runtime for state-of-the-art agent research, local inference, sandboxed tools, durable memory, and LAN Mesh resource sharing.
SSD-streaming inference engine for giant MoE models (Rust + CUDA). GLM 5.2 743B at 2 tok/s and Hy3 295B at 7 tok/s on two consumer 16GB GPUs. Zero-config multi-GPU: measures PCIe bandwidth, places attention and hot experts where they fit.
Aithy is a private local AI runtime for state-of-the-art agent research, local inference, sandboxed tools, durable memory, and LAN Mesh resource sharing.
Run GLM-5.2 (744B MoE) on a 25GB-RAM consumer machine — pure C, zero deps, experts streamed from disk. Tiny engine, immense model. 🐦
DeepSeek 4 Flash and PRO local inference engine for Metal, CUDA and ROCm
An occasional email when notable AI dev tools and models land in the directory. No spam, unsubscribe anytime.