Aithy
Aithy is a private local AI runtime for state-of-the-art agent research, local inference, sandboxed tools, durable memory, and LAN Mesh resource sharing.
25 tools in Local Inference.
Aithy is a private local AI runtime for state-of-the-art agent research, local inference, sandboxed tools, durable memory, and LAN Mesh resource sharing.
Run GLM-5.2 (744B MoE) on a 25GB-RAM consumer machine — pure C, zero deps, experts streamed from disk. Tiny engine, immense model. 🐦
DeepSeek 4 Flash and PRO local inference engine for Metal, CUDA and ROCm
Run frontier AI locally.
FastMLX is a high performance production ready API to host MLX models.
GPT4All: Run Local LLMs on Any Device. Open-source and available for commercial use.
Jan is an open source alternative to ChatGPT that runs 100% offline on your computer.
Local AI for Apple Silicon. Chat with open LLMs (Qwen, Gemma, DeepSeek, Mistral, Llama…) fully offline on your Mac via Apple MLX — private by design, zero telemetry.
🔥 The fastest local AI engine for Apple Silicon. Optimised for agentic use.
LLM inference in C/C++
Run local AI models like gpt-oss, Llama, Gemma, Qwen, and DeepSeek privately on your computer.
mac code — Claude Code, but it runs on your Mac for free. 35B AI agent at 30 tok/s via Apple Silicon flash-paging. $0/month.
Universal LLM Deployment Engine with ML Compilation
Up to 3× faster LLM decoding on Apple Silicon, lossless. Native MLX port of DeepSeek's DSpark & z-lab's DFlash speculative decoding — Gemma-4, Qwen3, Ornith-1.0, ternary Bonsai-27B.
Run LLMs with MLX
Quantize, fine-tune and serve LLMs locally on Apple Silicon (M1 to M5). MLX-native, no PyTorch, no cloud. On PyPI.
MLX-VLM is a package for inference and fine-tuning of Vision Language Models (VLMs) on your Mac using MLX.
A native macOS app for local LLMs. Native MTP decoding on Apple Silicon. Twice the speed, still exact.
Local AI, native to your Mac. Chat, serve, monitor, and connect MLX models from one macOS app.
Get up and running with Kimi-K2.6, GLM-5.2, MiniMax, DeepSeek, gpt-oss, Qwen, Gemma and other models.
Native macOS inference server built on MLX. Paged SSD KV caching, continuous batching, and drop-in API for Claude Code, OpenClaw, and Cursor.
Pure Rust + CUDA LLM inference engine — no PyTorch, OpenAI-compatible, serves Qwen3 to Kimi-K2
SSD-streaming inference engine for giant MoE models (Rust + CUDA). GLM 5.2 743B at 2 tok/s and Hy3 295B at 7 tok/s on two consumer 16GB GPUs. Zero-config multi-GPU: measures PCIe bandwidth, places attention and hot experts where they fit.
High-performance MLX-based LLM inference engine for macOS with native Swift implementation
Local UI to run and train LLMs and diffusion models, including Kimi K3, MiniMax-H3, Gemma 4, Qwen3.6, DeepSeek-V4, FLUX and more.
An occasional email when notable AI dev tools and models land in the directory. No spam, unsubscribe anytime.