AI models

The frontier and open-weights models developers are building on — with the specs that actually matter: context window, modalities, licensing and price per token.

Needle 3

Cactus Compute

A 121M-parameter open model for tool calling, structured extraction and text embedding on tiny devices. One set of weights ships as an intelligence ladder: every depth from 2 to 20 layers is a deployable model, an 8-29 MB CQ2-bit binary.

Open weights8K ctx121MTools

DeepSeek V4.1 Flash

DeepSeek

We introduce DeepSeek-V4.1-Flash, a multimodal Mixture-of-Experts (MoE) model with 552B backbone parameters and support for contexts of up to one million tokens. The model natively processes images and text, and generates text autoregressively.

Open weights1M ctx · 384K outReasoningToolsVision

GPT-6 Astra

OpenAI

Our most capable model, built for the hardest end-to-end work.

API1.05M ctx · 128K outReasoningToolsVision
$10 / $50 per 1M
24d ago

MAI-Transcribe-2

Microsoft AI

Turn noisy audio into precise, domain-specific transcripts, with leading FLEURS and Artificial Analysis accuracy scores.

APIAudio
24d ago

Gemini 3.8 Flash

Google

Our most intelligent workhorse model, delivering significant improvements from 3.7 Flash across software engineering, agentic tasks, and critical, multi-step reasoning in specialized domains.

API1M ctx · 64K outReasoningToolsVisionAudioPDF
$0.75 / $3.75 per 1M
25d ago

Gemini 3.8 Flash Cyber

Google

Our most capable cybersecurity model with frontier-level performance in vulnerability detection and automated patching, available to trusted defenders through our new Fairwind Program.

API
25d ago

Raptor 0.5

Osaurus

An 8B mixture-of-experts that activates only ~1B parameters per token, runs the full Osaurus tool surface in 6.3 GB, and stays resident on the 8–16 GB Macs most people own.

Open weights128K ctx7.9B total, ~1B active (MoE)ReasoningTools

Atlas

World Labs

Atlas is an omni model that we pretrained from scratch to natively operate on text, images, video, and 3D. It is a multimodal autoregressive diffusion transformer: all inputs are combined into a shared spatial context.

APIVisionImage outVideo out
26d ago

Claude Fable 5.1

Anthropic

Claude Fable 5.1 sets a new standard for coding, knowledge work, and long-running problem-solving tasks.

API1M ctx · 128K outReasoningToolsVisionPDF
$10 / $50 per 1M
26d ago

Claude Mythos 5.1

Anthropic

Mythos 5.1 is available only through our trusted access programs; its safeguards are specifically designed to support work in cybersecurity and the life sciences.

API1M ctx · 128K outReasoningToolsVisionPDF
$10 / $50 per 1M
26d ago

Gemini Omni 1.1 Flash

Google

Omni now delivers studio-quality video production, including the ability to extend a scene, first and last frame interpolation, crisp 4K upscaling, faster prototyping, and more.

APIVisionVideo out
4w ago

Gemini 3.5 Transcribe

Google

Our most precise speech-to-text model yet, designed for intelligent voice interactions.

APIToolsAudio
$2 / $12 per 1M
5w ago

Gemini 3.5 Transcribe Live

Google

Our most precise speech-to-text model yet, designed for intelligent voice interactions.

APIToolsAudio
$3.5 / $21 per 1M
5w ago

GLM-5.3-Flash

Zhipu AI

GLM-5.3-Flash is the first native multimodal model in the GLM-5 series, delivering stronger intelligence than GLM-5.2 while maintaining an exceptionally cost-efficient architecture.

Open weights1M ctx · 128K out320B total, 18B activeReasoningToolsVisionPDF

GLiNER2.5 Base

Fastino Labs

A new boundary-prediction architecture replaces span enumeration, adding joint entity-relation extraction, constrained classification, and unlimited span length.

Open weights

GLiNER2.5 Multilingual

Fastino Labs

A new boundary-prediction architecture replaces span enumeration, adding joint entity-relation extraction, constrained classification, and unlimited span length.

Open weights

GLiNER2.5 Small

Fastino Labs

A new boundary-prediction architecture replaces span enumeration, adding joint entity-relation extraction, constrained classification, and unlimited span length.

Open weights

Qwen3.8 Flash Next

Qwen

An open-weights multimodal MoE model that doubles as an early preview of the Qwen4 architecture, the same role Qwen3-Next played for Qwen3.5. It pairs Gated DeltaNet with Qwen Sparse Attention, widens the residual stream into four gated branches, and adds 51B N-gram embedding parameters that cost almost nothing per token. Natively 262K context, extensible to 1M with YaRN.

Open weights256K ctx125B total, 6B active, plus 51B N-gram embeddingsReasoningToolsVision

GLM-5.3

Zhipu AI

GLM-5.3 is the most capable open-weights model for coding, with a 50% improvement over GLM-5.2 on our in-house Z.ai Code Bench.

API1M ctx · 128K outReasoningTools
$1.4 / $4.4 per 1M
6w ago

Qwen3.8 27B

Qwen

Built on the architectural foundation of Qwen3.5, Qwen3.8 delivers substantial gains across coding, professional work, research, and long-horizon agentic tasks. Qwen3.8-27B brings these advances to a compact, deployment-friendly dense model: a native vision-language model that understands images and videos, with flexible thinking control, designed to carry complex, multi-step tasks through to completion with greater reliability.

Open weights256K ctx · 128K out27B denseReasoningToolsVision

Gemini 3.7 Flash

Google

Our most intelligent workhorse model yet for coding and agents.

API1M ctx · 64K outReasoningToolsVisionAudioPDF
$0.75 / $3.75 per 1M
6w ago

Grok 4.6

xAI

xAI's frontier Grok model, tuned for long-running agents, coding, knowledge work, and visual projects.

API500K ctx · 500K outReasoningToolsVisionPDF
$2 / $6 per 1M
7w ago

LFM2.5-VL-3B

Liquid AI

LFM2.5-VL-3B is the multimodal variant of LFM2.5, pairing the LFM2.5-2.6B backbone with a SigLIP2 NaFlex 400M vision encoder. It answers directly rather than reasoning, and adds screen understanding, grounding and function calling for on-device use.

Open weights32K ctx3.1BToolsVision

S1-mini

Superwhisper

A 0.6B-parameter text normalizer for speech-to-text output. It takes a raw ASR transcript and rewrites it as clean written text: fillers removed, false starts and self-corrections resolved to the value the speaker landed on, punctuation and capitalization applied, and spoken numbers, dates, times, currency and email addresses rendered in written form.

Open weights40K ctx596M

LTX-2.5

LTX

LTX-2.5 generates multi-shot scenes in one pass, edits real footage, and exports cinema-grade EXR. Open weights you can fine-tune and run on your hardware.

Open weightsVisionAudioVideo outAudio out

Needle 2

Cactus Compute

An open 45M-parameter model for tool calling, device use, and structured extraction. Needle 2 runs as a 14 MB binary in 28 MB of session RAM.

Open weights256 ctx45MTools

Nemotron 3.5 Lightning

NVIDIA

A customizable open 30B MoE model with 3B active parameters, providing optimal high-volume execution for autonomous agents.

Open weights256K ctx · 256K out30B total, 3B active (MoE)ReasoningTools

Ling 3.0 Tiny

inclusionAI

We are introducing Ling-3.0-tiny, a lightweight hybrid reasoning MoE model with 7.9B total parameters and only 1.3B activated parameters per token. It is designed to deliver strong reasoning and agentic capabilities at low inference cost, making advanced model capabilities more accessible for local and resource-constrained deployment.

Open weights256K ctx · 32K out7.9B total, 1.3B active (MoE)ReasoningTools

Muse Glimmer 30B

Meta

Muse Glimmer is a 30-billion-parameter causal language model with a dedicated perception encoder, distilled from Muse Spark and purpose-built for autonomous agentic tasks on consumer hardware. The model integrates multi-step reasoning, reliable tool use, multimodal understanding, and failure recovery into a single model that runs locally without requiring cloud infrastructure or network access.

Open weights128K ctx29.6BReasoningToolsVision

TwiL-LM 1.7B

webAI

A LoRA adapter over SmolLM2-1.7B-Instruct for formal logic — FOL translation, entailment, semantic parsing and Lean assistance — quantized to a 1.06 GB download that runs on a phone.

Open weights8K ctx1.7B

TwiL-LM3

webAI

A 3B formal-logic reasoning model built from SmolLM3-3B by LoRA fine-tuning, checkpoint fusion, WiSE-FT interpolation and entropy-weighted GRPO, which beats gpt-oss-120b on four of five lanes of webAI's formal-reasoning suite.

Open weights64K ctx3BReasoning

FLUX 3

Black Forest Labs

FLUX 3 is our new multimodal foundation model. It jointly learns from images, videos, and audio within a unified architecture, because what it needs to learn is not any one of these elements in isolation. Instead, a model must learn a representation of the world: how objects hold together, how things move, and how events sound.

APIVisionAudioVideo outAudio out
8w ago

LFM2.5-2.6B

Liquid AI

LFM2.5-2.6B is part of LFM2.5, a family of hybrid models designed for on-device deployment. It builds on the LFM2 architecture with a 128K context window and agentic post-training.

Open weights128K ctx2.6BTools

Shieldstral

Mistral AI

A 3B open-weights, policy-adaptive multimodal safety classifier that matches models up to 7x its size on text safety and sets a new state of the art on multimodal moderation.

Open weights256K ctx3BVision

MiniMax H3

MiniMax

MiniMax H3 is a general-purpose, omni-modal generative system. It supports unified understanding of multimodal contexts composed of text, images, video, and audio, and can generate video with native stereo audio at resolutions up to 2K and durations of up to 15 seconds.

Open weights33BVisionAudioVideo outAudio out

Inkling-Small

Thinking Machines Lab

A quarter the size of Inkling at comparable performance: a multimodal Mixture-of-Experts reasoner (276B total, 12B active) with controllable reasoning effort, fine-tunable on Tinker.

Open weights

Claude Opus 5

Anthropic

Opus 5 is a step change improvement for the Opus tier powering long-running agents while delivering improvements in coding and professional work.

API1M ctx · 128K outReasoningToolsVisionPDF
$5 / $25 per 1M
9w ago

Gemini 3.5 Flash-Lite

Google

Our fastest, most cost-effective 3.5-class model, delivering 350 output tokens per second.

API1M ctx · 64K outReasoningToolsVisionAudioPDF
$0.3 / $2.5 per 1M
10w ago

Gemini 3.6 Flash

Google

Our workhorse model that delivers better coding, knowledge work, and multimodal performance.

API1M ctx · 64K outReasoningToolsVisionAudioPDF
$0.75 / $3.75 per 1M
10w ago

Laguna S 2.1

Poolside

The most capable agentic coding model in its weight class by a wide margin.

Open weights1M ctx · 32K out118B total, ~8B active (MoE)ReasoningTools

Kimi K3

Moonshot AI

Kimi K3 is a 2.8T-parameter model built on Kimi Delta Attention and Attention Residuals, with native vision capabilities and a 1-million-token context window — the world's first open 3T-class model.

Open weights1M ctx · 128K outReasoningToolsVision

Inkling

Thinking Machines Lab

Thinking Machines Lab's first open-weights model: a multimodal Mixture-of-Experts reasoner (975B total, 41B active) with controllable reasoning effort, fine-tunable on Tinker.

Open weights64K ctx · 64K outReasoningToolsVision

Bonsai 27B

PrismML

The first 27B-class model to run on a phone, based on Qwen3.6 27B: a multimodal flagship shipping in ternary (5.9 GB) and 1-bit (3.9 GB) forms with speculative decoding.

Open weights256K ctxReasoningToolsVision
Apache-2.0
11w ago

GPT-5.6 Luna

OpenAI

GPT-5.6 model optimized for cost-sensitive workloads.

API1.05M ctx · 128K outReasoningToolsVisionPDF
$0.2 / $1.2 per 1M
11w ago

GPT-5.6 Sol

OpenAI

Frontier model for complex professional work.

API1.05M ctx · 128K outReasoningToolsVisionPDF
$4 / $20 per 1M
11w ago

GPT-5.6 Terra

OpenAI

GPT-5.6 model that balances intelligence and cost.

API1.05M ctx · 128K outReasoningToolsVisionPDF
$2 / $12 per 1M
11w ago

Grok 4.5

xAI

xAI's Grok 4.5 flagship for chat, coding, and agentic tool use, with lower hallucination risk.

API500K ctx · 500K outReasoningToolsVisionPDF
$2 / $6 per 1M
12w ago

SWE-1.7

Cognition

The most capable model Cognition has trained so far. It reaches frontier-level intelligence at a much lower cost, advancing the cost-performance Pareto curve.

APITools
12w ago

Claude Sonnet 5

Anthropic

Claude Sonnet 5 is built to be the most agentic Sonnet model yet. It can make plans, use tools like browsers and terminals, and run autonomously at a level that, just a few months ago, required larger and more expensive models.

API1M ctx · 128K outReasoningToolsVisionPDF
$2 / $10 per 1M
29 Jun 2026

GLM-5.2

Zhipu AI

We're introducing GLM-5.2, our latest flagship model for long-horizon tasks.

Open weights1M ctx · 128K outReasoningTools

DeepSeek V4 Flash

DeepSeek

DeepSeek-V4-Flash-0731 is the official release of DeepSeek-V4-Flash, superseding the preview version, with substantially enhanced agentic capabilities.

Open weights1M ctx · 384K outReasoningTools
↓ 4.4M♥ 3.9K
24 Apr 2026

DeepSeek V4 Pro

DeepSeek

DeepSeek-V4-Pro is a 1.6T-parameter Mixture-of-Experts model with 49B activated, supporting a one-million-token context. The 0813 revision is its GA release.

Open weights1M ctx · 384K out1.6T total, 49B active (MoE)ReasoningTools

GPT-5.5

OpenAI

A new class of intelligence for coding and professional work.

API1.05M ctx · 128K outReasoningToolsVisionPDF
$5 / $30 per 1M
23 Apr 2026

Kimi K2.6

Moonshot AI

Kimi K2.6 is an open-source, native multimodal agentic model that advances practical capabilities in long-horizon coding, coding-driven design, proactive autonomous execution, and swarm-based task orchestration.

Open weights256K ctx · 256K out1TReasoningToolsVision

Gemini 3.5 Flash-Cyber

Google

Fine-tuned for finding and fixing cybersecurity vulnerabilities at a lower price per token than larger models.

API

Keep up with the tools

An occasional email when notable AI dev tools and models land in the directory. No spam, unsubscribe anytime.