articleAI Local4 Sept 2026

I’m running an Opus-level coding agent locally at nearly 2x Claude Opus speed, for free.

Nov 2025 was a tipping point for AI coding with commercial frontier models. Aug 2026 is the beginning of the same thing for local AI. Pretty darned exciting.

Julian Harris
ailocal.substack.com
articleLiquid AI20 Aug 2026

LFM2.5-DSpark: Up to 3.2x Faster Inference from H100 to MacBook

LFM2.5-DSpark speculative decoding delivers up to 3.2x faster inference on GPU and 2.9x on-device. Day-one llama.cpp and SGLang support.

Liquid AI
liquid.ai
articleX25 May 2026

Every llama-server Flag Explained: The Tuning Guide For Local LLMs

witcheer@witcheer
x.com
articleX20 May 2026

llama.cpp - Run Local LLMs On Your GPU

witcheer@witcheer
x.com

More in Local Inference

View all tools
A

Aithy

Aithy is a private local AI runtime for state-of-the-art agent research, local inference, sandboxed tools, durable memory, and LAN Mesh resource sharing.

Local Inference
Atomic Chat logo

Atomic Chat

Local AI app and inference engine for agents. Run open-weight LLMs locally — private, 100% offline on your computer.

Local Inference
C

Colibri

Run GLM-5.2 (744B MoE) on a 25GB-RAM consumer machine — pure C, zero deps, experts streamed from disk. Tiny engine, immense model. 🐦

Local Inference

Keep up with the tools

An occasional email when notable AI dev tools and models land in the directory. No spam, unsubscribe anytime.