Local Inference

25 tools in Local Inference.

A

Aithy

Aithy is a private local AI runtime for state-of-the-art agent research, local inference, sandboxed tools, durable memory, and LAN Mesh resource sharing.

Local Inference
C

Colibri

Run GLM-5.2 (744B MoE) on a 25GB-RAM consumer machine — pure C, zero deps, experts streamed from disk. Tiny engine, immense model. 🐦

Local Inference
D

DwarfStar

DeepSeek 4 Flash and PRO local inference engine for Metal, CUDA and ROCm

Local Inference
exo logo

exo

Run frontier AI locally.

Local Inference
FastMLX logo

FastMLX

FastMLX is a high performance production ready API to host MLX models.

Local Inference
GPT4All logo

GPT4All

GPT4All: Run Local LLMs on Any Device. Open-source and available for commercial use.

Local Inference
Jan logo

Jan

Jan is an open source alternative to ChatGPT that runs 100% offline on your computer.

Local Inference
Ka1zen logo

Ka1zen

Local AI for Apple Silicon. Chat with open LLMs (Qwen, Gemma, DeepSeek, Mistral, Llama…) fully offline on your Mac via Apple MLX — private by design, zero telemetry.

Local InferenceDesktop Applications
L

Lightning-MLX

🔥 The fastest local AI engine for Apple Silicon. Optimised for agentic use.

Local Inference
llama.cpp logo

llama.cpp

LLM inference in C/C++

Local Inference
LM Studio logo

LM Studio

Run local AI models like gpt-oss, Llama, Gemma, Qwen, and DeepSeek privately on your computer.

Local Inference
M

mac-code

mac code — Claude Code, but it runs on your Mac for free. 35B AI agent at 30 tok/s via Apple Silicon flash-paging. $0/month.

Local Inference
MLC LLM logo

MLC LLM

Universal LLM Deployment Engine with ML Compilation

Local Inference
M

mlx-dspark

Up to 3× faster LLM decoding on Apple Silicon, lossless. Native MLX port of DeepSeek's DSpark & z-lab's DFlash speculative decoding — Gemma-4, Qwen3, Ornith-1.0, ternary Bonsai-27B.

Local Inference
mlx-lm logo

mlx-lm

Run LLMs with MLX

Local Inference
mlx-optiq logo

mlx-optiq

Quantize, fine-tune and serve LLMs locally on Apple Silicon (M1 to M5). MLX-native, no PyTorch, no cloud. On PyPI.

Local Inference
M

MLX-VLM

MLX-VLM is a package for inference and fine-tuning of Vision Language Models (VLMs) on your Mac using MLX.

Local Inference
MTPLX logo

MTPLX

A native macOS app for local LLMs. Native MTP decoding on Apple Silicon. Twice the speed, still exact.

Local Inference
Nativ logo

Nativ

Local AI, native to your Mac. Chat, serve, monitor, and connect MLX models from one macOS app.

Local Inference
Ollama logo

Ollama

Get up and running with Kimi-K2.6, GLM-5.2, MiniMax, DeepSeek, gpt-oss, Qwen, Gemma and other models.

Local Inference
oMLX logo

oMLX

Native macOS inference server built on MLX. Paged SSD KV caching, continuous batching, and drop-in API for Claude Code, OpenClaw, and Cursor.

Local Inference
openinfer logo

openinfer

Pure Rust + CUDA LLM inference engine — no PyTorch, OpenAI-compatible, serves Qwen3 to Kimi-K2

Local Inference
P

pulsar

SSD-streaming inference engine for giant MoE models (Rust + CUDA). GLM 5.2 743B at 2 tok/s and Hy3 295B at 7 tok/s on two consumer 16GB GPUs. Zero-config multi-GPU: measures PCIe bandwidth, places attention and hot experts where they fit.

Local Inference
Swama logo

Swama

High-performance MLX-based LLM inference engine for macOS with native Swift implementation

Local Inference
Unsloth logo

Unsloth

Local UI to run and train LLMs and diffusion models, including Kimi K3, MiniMax-H3, Gemma 4, Qwen3.6, DeepSeek-V4, FLUX and more.

Local InferenceFine Tuning

Keep up with the tools

An occasional email when notable AI dev tools and models land in the directory. No spam, unsubscribe anytime.