mlx-serve logo

mlx-serve

Native LLM inference server for Apple Silicon. OpenAI + Anthropic API compatible. No Python. Includes MLX Core macOS app with chat, agent mode, and tool calling.

Categories

blogHuman Learning27 Aug 2026

I Tried Open Weight Models to Cut My AI Bill to Zero

Your AI product is an overnight success BUT your AI bills is giving you a heartache.

Aishwarya Raghavan
humanlearning.substack.com

More in Local Inference

View all tools
A

Aithy

Aithy is a private local AI runtime for state-of-the-art agent research, local inference, sandboxed tools, durable memory, and LAN Mesh resource sharing.

Local Inference
Atomic Chat logo

Atomic Chat

Local AI app and inference engine for agents. Run open-weight LLMs locally — private, 100% offline on your computer.

Local Inference
C

Colibri

Run GLM-5.2 (744B MoE) on a 25GB-RAM consumer machine — pure C, zero deps, experts streamed from disk. Tiny engine, immense model. 🐦

Local Inference

Keep up with the tools

An occasional email when notable AI dev tools and models land in the directory. No spam, unsubscribe anytime.