oMLX logo

oMLX

Native macOS inference server built on MLX. Paged SSD KV caching, continuous batching, and drop-in API for Claude Code, OpenClaw, and Cursor.

Categories

blogJoshua Warren27 Aug 2026

Qwen3.8 is running on the Apple Neural Engine under Omarchy Linux

An M1 MacBook Pro running Omarchy Linux now sends real Qwen3.8 weights through the Apple Neural Engine. The full 24-layer model path works with persistent DeltaNet and attention state. A macOS control shows where the Linux runtime needs to go next.

Joshua Warren
joshuawarren.com

More in Local Inference

View all tools
A

Aithy

Aithy is a private local AI runtime for state-of-the-art agent research, local inference, sandboxed tools, durable memory, and LAN Mesh resource sharing.

Local Inference
Atomic Chat logo

Atomic Chat

Local AI app and inference engine for agents. Run open-weight LLMs locally — private, 100% offline on your computer.

Local Inference
C

Colibri

Run GLM-5.2 (744B MoE) on a 25GB-RAM consumer machine — pure C, zero deps, experts streamed from disk. Tiny engine, immense model. 🐦

Local Inference

Keep up with the tools

An occasional email when notable AI dev tools and models land in the directory. No spam, unsubscribe anytime.