A small Metal inference server for one checkpoint: Qwen3.6-35B-A3B converted to MLX affine 4-bit weights. Lily exposes a minimal subset of the OpenAI chat completions API and always decodes greedily.

Categories

blogPerplexity AI1 Sept 2026

Optimizing On-Device Inference for Apple Silicon

Perplexity's engineering team details Lily, a Rust-and-Metal local inference engine built specifically for Apple silicon and Qwen3.6-35B-A3B, which averages 1.23x MLX-LM's prefill and 1.35x its decode throughput on an M5 Max.

Perplexity Engineering
perplexity.ai

More in Local Inference

View all tools
A

Aithy

Aithy is a private local AI runtime for state-of-the-art agent research, local inference, sandboxed tools, durable memory, and LAN Mesh resource sharing.

Local Inference
Atomic Chat logo

Atomic Chat

Local AI app and inference engine for agents. Run open-weight LLMs locally — private, 100% offline on your computer.

Local Inference
C

Colibri

Run GLM-5.2 (744B MoE) on a 25GB-RAM consumer machine — pure C, zero deps, experts streamed from disk. Tiny engine, immense model. 🐦

Local Inference

Keep up with the tools

An occasional email when notable AI dev tools and models land in the directory. No spam, unsubscribe anytime.