The frontier and open-weights models developers are building on — with the specs that actually matter: context window, modalities, licensing and price per token.
Needle 3
Cactus Compute
A 121M-parameter open model for tool calling, structured extraction and text embedding on tiny devices. One set of weights ships as an intelligence ladder: every depth from 2 to 20 layers is a deployable model, an 8-29 MB CQ2-bit binary.
We introduce DeepSeek-V4.1-Flash, a multimodal Mixture-of-Experts (MoE) model with 552B backbone parameters and support for contexts of up to one million tokens. The model natively processes images and text, and generates text autoregressively.
Our most capable model, built for the hardest end-to-end work.
API1.05M ctx · 128K outReasoningToolsVision
$10 / $50 per 1M
24d ago
MAI-Transcribe-2
Microsoft AI
Turn noisy audio into precise, domain-specific transcripts, with leading FLEURS and Artificial Analysis accuracy scores.
APIAudio
24d ago
Gemini 3.8 Flash
Google
Our most intelligent workhorse model, delivering significant improvements from 3.7 Flash across software engineering, agentic tasks, and critical, multi-step reasoning in specialized domains.
API1M ctx · 64K outReasoningToolsVisionAudioPDF
$0.75 / $3.75 per 1M
25d ago
Gemini 3.8 Flash Cyber
Google
Our most capable cybersecurity model with frontier-level performance in vulnerability detection and automated patching, available to trusted defenders through our new Fairwind Program.
API
25d ago
Raptor 0.5
Osaurus
An 8B mixture-of-experts that activates only ~1B parameters per token, runs the full Osaurus tool surface in 6.3 GB, and stays resident on the 8–16 GB Macs most people own.
Open weights128K ctx7.9B total, ~1B active (MoE)ReasoningTools
Atlas is an omni model that we pretrained from scratch to natively operate on text, images, video, and 3D. It is a multimodal autoregressive diffusion transformer: all inputs are combined into a shared spatial context.
APIVisionImage outVideo out
26d ago
Claude Fable 5.1
Anthropic
Claude Fable 5.1 sets a new standard for coding, knowledge work, and long-running problem-solving tasks.
API1M ctx · 128K outReasoningToolsVisionPDF
$10 / $50 per 1M
26d ago
Claude Mythos 5.1
Anthropic
Mythos 5.1 is available only through our trusted access programs; its safeguards are specifically designed to support work in cybersecurity and the life sciences.
API1M ctx · 128K outReasoningToolsVisionPDF
$10 / $50 per 1M
26d ago
Gemini Omni 1.1 Flash
Google
Omni now delivers studio-quality video production, including the ability to extend a scene, first and last frame interpolation, crisp 4K upscaling, faster prototyping, and more.
APIVisionVideo out
4w ago
Gemini 3.5 Transcribe
Google
Our most precise speech-to-text model yet, designed for intelligent voice interactions.
APIToolsAudio
$2 / $12 per 1M
5w ago
Gemini 3.5 Transcribe Live
Google
Our most precise speech-to-text model yet, designed for intelligent voice interactions.
APIToolsAudio
$3.5 / $21 per 1M
5w ago
GLM-5.3-Flash
Zhipu AI
GLM-5.3-Flash is the first native multimodal model in the GLM-5 series, delivering stronger intelligence than GLM-5.2 while maintaining an exceptionally cost-efficient architecture.
Open weights1M ctx · 128K out320B total, 18B activeReasoningToolsVisionPDF
An open-weights multimodal MoE model that doubles as an early preview of the Qwen4 architecture, the same role Qwen3-Next played for Qwen3.5. It pairs Gated DeltaNet with Qwen Sparse Attention, widens the residual stream into four gated branches, and adds 51B N-gram embedding parameters that cost almost nothing per token. Natively 262K context, extensible to 1M with YaRN.
Open weights256K ctx125B total, 6B active, plus 51B N-gram embeddingsReasoningToolsVision
GLM-5.3 is the most capable open-weights model for coding, with a 50% improvement over GLM-5.2 on our in-house Z.ai Code Bench.
API1M ctx · 128K outReasoningTools
$1.4 / $4.4 per 1M
6w ago
Qwen3.8 27B
Qwen
Built on the architectural foundation of Qwen3.5, Qwen3.8 delivers substantial gains across coding, professional work, research, and long-horizon agentic tasks. Qwen3.8-27B brings these advances to a compact, deployment-friendly dense model: a native vision-language model that understands images and videos, with flexible thinking control, designed to carry complex, multi-step tasks through to completion with greater reliability.
Open weights256K ctx · 128K out27B denseReasoningToolsVision
Our most intelligent workhorse model yet for coding and agents.
API1M ctx · 64K outReasoningToolsVisionAudioPDF
$0.75 / $3.75 per 1M
6w ago
Grok 4.6
xAI
xAI's frontier Grok model, tuned for long-running agents, coding, knowledge work, and visual projects.
API500K ctx · 500K outReasoningToolsVisionPDF
$2 / $6 per 1M
7w ago
LFM2.5-VL-3B
Liquid AI
LFM2.5-VL-3B is the multimodal variant of LFM2.5, pairing the LFM2.5-2.6B backbone with a SigLIP2 NaFlex 400M vision encoder. It answers directly rather than reasoning, and adds screen understanding, grounding and function calling for on-device use.
A 0.6B-parameter text normalizer for speech-to-text output. It takes a raw ASR transcript and rewrites it as clean written text: fillers removed, false starts and self-corrections resolved to the value the speaker landed on, punctuation and capitalization applied, and spoken numbers, dates, times, currency and email addresses rendered in written form.
LTX-2.5 generates multi-shot scenes in one pass, edits real footage, and exports cinema-grade EXR. Open weights you can fine-tune and run on your hardware.
We are introducing Ling-3.0-tiny, a lightweight hybrid reasoning MoE model with 7.9B total parameters and only 1.3B activated parameters per token. It is designed to deliver strong reasoning and agentic capabilities at low inference cost, making advanced model capabilities more accessible for local and resource-constrained deployment.
Open weights256K ctx · 32K out7.9B total, 1.3B active (MoE)ReasoningTools
Muse Glimmer is a 30-billion-parameter causal language model with a dedicated perception encoder, distilled from Muse Spark and purpose-built for autonomous agentic tasks on consumer hardware. The model integrates multi-step reasoning, reliable tool use, multimodal understanding, and failure recovery into a single model that runs locally without requiring cloud infrastructure or network access.
A LoRA adapter over SmolLM2-1.7B-Instruct for formal logic — FOL translation, entailment, semantic parsing and Lean assistance — quantized to a 1.06 GB download that runs on a phone.
A 3B formal-logic reasoning model built from SmolLM3-3B by LoRA fine-tuning, checkpoint fusion, WiSE-FT interpolation and entropy-weighted GRPO, which beats gpt-oss-120b on four of five lanes of webAI's formal-reasoning suite.
FLUX 3 is our new multimodal foundation model. It jointly learns from images, videos, and audio within a unified architecture, because what it needs to learn is not any one of these elements in isolation. Instead, a model must learn a representation of the world: how objects hold together, how things move, and how events sound.
APIVisionAudioVideo outAudio out
8w ago
LFM2.5-2.6B
Liquid AI
LFM2.5-2.6B is part of LFM2.5, a family of hybrid models designed for on-device deployment. It builds on the LFM2 architecture with a 128K context window and agentic post-training.
A 3B open-weights, policy-adaptive multimodal safety classifier that matches models up to 7x its size on text safety and sets a new state of the art on multimodal moderation.
MiniMax H3 is a general-purpose, omni-modal generative system. It supports unified understanding of multimodal contexts composed of text, images, video, and audio, and can generate video with native stereo audio at resolutions up to 2K and durations of up to 15 seconds.
A quarter the size of Inkling at comparable performance: a multimodal Mixture-of-Experts reasoner (276B total, 12B active) with controllable reasoning effort, fine-tunable on Tinker.
Kimi K3 is a 2.8T-parameter model built on Kimi Delta Attention and Attention Residuals, with native vision capabilities and a 1-million-token context window — the world's first open 3T-class model.
The first 27B-class model to run on a phone, based on Qwen3.6 27B: a multimodal flagship shipping in ternary (5.9 GB) and 1-bit (3.9 GB) forms with speculative decoding.
Open weights256K ctxReasoningToolsVision
Apache-2.0
11w ago
GPT-5.6 Luna
OpenAI
GPT-5.6 model optimized for cost-sensitive workloads.
API1.05M ctx · 128K outReasoningToolsVisionPDF
$0.2 / $1.2 per 1M
11w ago
GPT-5.6 Sol
OpenAI
Frontier model for complex professional work.
API1.05M ctx · 128K outReasoningToolsVisionPDF
$4 / $20 per 1M
11w ago
GPT-5.6 Terra
OpenAI
GPT-5.6 model that balances intelligence and cost.
API1.05M ctx · 128K outReasoningToolsVisionPDF
$2 / $12 per 1M
11w ago
Grok 4.5
xAI
xAI's Grok 4.5 flagship for chat, coding, and agentic tool use, with lower hallucination risk.
API500K ctx · 500K outReasoningToolsVisionPDF
$2 / $6 per 1M
12w ago
SWE-1.7
Cognition
The most capable model Cognition has trained so far. It reaches frontier-level intelligence at a much lower cost, advancing the cost-performance Pareto curve.
APITools
12w ago
Claude Sonnet 5
Anthropic
Claude Sonnet 5 is built to be the most agentic Sonnet model yet. It can make plans, use tools like browsers and terminals, and run autonomously at a level that, just a few months ago, required larger and more expensive models.
API1M ctx · 128K outReasoningToolsVisionPDF
$2 / $10 per 1M
29 Jun 2026
GLM-5.2
Zhipu AI
We're introducing GLM-5.2, our latest flagship model for long-horizon tasks.
DeepSeek-V4-Flash-0731 is the official release of DeepSeek-V4-Flash, superseding the preview version, with substantially enhanced agentic capabilities.
DeepSeek-V4-Pro is a 1.6T-parameter Mixture-of-Experts model with 49B activated, supporting a one-million-token context. The 0813 revision is its GA release.
Open weights1M ctx · 384K out1.6T total, 49B active (MoE)ReasoningTools
A new class of intelligence for coding and professional work.
API1.05M ctx · 128K outReasoningToolsVisionPDF
$5 / $30 per 1M
23 Apr 2026
Kimi K2.6
Moonshot AI
Kimi K2.6 is an open-source, native multimodal agentic model that advances practical capabilities in long-horizon coding, coding-driven design, proactive autonomous execution, and swarm-based task orchestration.
Open weights256K ctx · 256K out1TReasoningToolsVision