Grok 4.6
xAI
xAI's frontier Grok model, tuned for long-running agents, coding, knowledge work, and visual projects.
The frontier and open-weights models developers are building on — with the specs that actually matter: context window, modalities, licensing and price per token.
xAI
xAI's frontier Grok model, tuned for long-running agents, coding, knowledge work, and visual projects.
Liquid AI
LFM2.5-VL-3B is the multimodal variant of LFM2.5, pairing the LFM2.5-2.6B backbone with a SigLIP2 NaFlex 400M vision encoder. It answers directly rather than reasoning, and adds screen understanding, grounding and function calling for on-device use.
LTX
LTX-2.5 generates multi-shot scenes in one pass, edits real footage, and exports cinema-grade EXR. Open weights you can fine-tune and run on your hardware.
Cactus Compute
An open 45M-parameter model for tool calling, device use, and structured extraction. Needle 2 runs as a 14 MB binary in 28 MB of session RAM.
NVIDIA
A customizable open 30B MoE model with 3B active parameters, providing optimal high-volume execution for autonomous agents.
inclusionAI
We are introducing Ling-3.0-tiny, a lightweight hybrid reasoning MoE model with 7.9B total parameters and only 1.3B activated parameters per token. It is designed to deliver strong reasoning and agentic capabilities at low inference cost, making advanced model capabilities more accessible for local and resource-constrained deployment.
Meta
Muse Glimmer is a 30-billion-parameter causal language model with a dedicated perception encoder, distilled from Muse Spark and purpose-built for autonomous agentic tasks on consumer hardware. The model integrates multi-step reasoning, reliable tool use, multimodal understanding, and failure recovery into a single model that runs locally without requiring cloud infrastructure or network access.
webAI
A LoRA adapter over SmolLM2-1.7B-Instruct for formal logic — FOL translation, entailment, semantic parsing and Lean assistance — quantized to a 1.06 GB download that runs on a phone.
webAI
A 3B formal-logic reasoning model built from SmolLM3-3B by LoRA fine-tuning, checkpoint fusion, WiSE-FT interpolation and entropy-weighted GRPO, which beats gpt-oss-120b on four of five lanes of webAI's formal-reasoning suite.
Black Forest Labs
FLUX 3 is our new multimodal foundation model. It jointly learns from images, videos, and audio within a unified architecture, because what it needs to learn is not any one of these elements in isolation. Instead, a model must learn a representation of the world: how objects hold together, how things move, and how events sound.
Liquid AI
LFM2.5-2.6B is part of LFM2.5, a family of hybrid models designed for on-device deployment. It builds on the LFM2 architecture with a 128K context window and agentic post-training.
Mistral AI
A 3B open-weights, policy-adaptive multimodal safety classifier that matches models up to 7x its size on text safety and sets a new state of the art on multimodal moderation.
MiniMax
MiniMax H3 is a general-purpose, omni-modal generative system. It supports unified understanding of multimodal contexts composed of text, images, video, and audio, and can generate video with native stereo audio at resolutions up to 2K and durations of up to 15 seconds.
Thinking Machines Lab
A quarter the size of Inkling at comparable performance: a multimodal Mixture-of-Experts reasoner (276B total, 12B active) with controllable reasoning effort, fine-tunable on Tinker.
Anthropic
Opus 5 is a step change improvement for the Opus tier powering long-running agents while delivering improvements in coding and professional work.
Our fastest, most cost-effective 3.5-class model, delivering 350 output tokens per second.
Our workhorse model that delivers better coding, knowledge work, and multimodal performance.
Poolside
The most capable agentic coding model in its weight class by a wide margin.
Moonshot AI
Kimi K3 is a 2.8T-parameter model built on Kimi Delta Attention and Attention Residuals, with native vision capabilities and a 1-million-token context window — the world's first open 3T-class model.
Thinking Machines Lab
Thinking Machines Lab's first open-weights model: a multimodal Mixture-of-Experts reasoner (975B total, 41B active) with controllable reasoning effort, fine-tunable on Tinker.
PrismML
The first 27B-class model to run on a phone, based on Qwen3.6 27B: a multimodal flagship shipping in ternary (5.9 GB) and 1-bit (3.9 GB) forms with speculative decoding.
OpenAI
GPT-5.6 model optimized for cost-sensitive workloads.
OpenAI
Frontier model for complex professional work.
OpenAI
GPT-5.6 model that balances intelligence and cost.
xAI
xAI's Grok 4.5 flagship for chat, coding, and agentic tool use, with lower hallucination risk.
Cognition
The most capable model Cognition has trained so far. It reaches frontier-level intelligence at a much lower cost, advancing the cost-performance Pareto curve.
Anthropic
Claude Sonnet 5 is built to be the most agentic Sonnet model yet. It can make plans, use tools like browsers and terminals, and run autonomously at a level that, just a few months ago, required larger and more expensive models.
Zhipu AI
We're introducing GLM-5.2, our latest flagship model for long-horizon tasks.
DeepSeek
DeepSeek-V4-Flash-0731 is the official release of DeepSeek-V4-Flash, superseding the preview version, with substantially enhanced agentic capabilities.
DeepSeek
DeepSeek-V4-Pro is a 1.6T-parameter Mixture-of-Experts model with 49B activated, supporting a one-million-token context. The 0813 revision is its GA release.
OpenAI
A new class of intelligence for coding and professional work.
Moonshot AI
Kimi K2.6 is an open-source, native multimodal agentic model that advances practical capabilities in long-horizon coding, coding-driven design, proactive autonomous execution, and swarm-based task orchestration.
Fine-tuned for finding and fixing cybersecurity vulnerabilities at a lower price per token than larger models.
An occasional email when notable AI dev tools and models land in the directory. No spam, unsubscribe anytime.