Articles

Announcements, blog posts, tutorials and papers from across the AI-devtools ecosystem — linked to the tools and models they're about.

X

/show-me: compact visual representations for coding agents

tl;dr make your agent converse visually instead of in walls of prose. Lighter and faster than HTML, good enough for most dev-work shaped problems.

x.com
X

Agents Need an AST Layer

In the last two days I built two codemods. One helps migrate a codebase from one major version of a framework to the next.

x.com
Vercel

Building a software factory for AI SDK

We built a software factory that autonomously processes issues and PRs for the AI SDK, with humans in control of every merge. Four weeks in, it authors 25-40% of merged PRs.

Lars Grammel and Eric DoddsAI SDK
vercel.com
X

Grok 4.6 – A field guide

Grok 4.6 is out! I've used it for a few weeks as my daily driver across the normal mix of coding and knowledge work, and built a few projects with it specifically to push on where it holds up.

eric zakariasson@ericzakariassonGrok 4.6
x.com
tailscale.com

How Tailscale helped find the SQLite WAL-Reset bug

Tailscale and SQLite developers traced maddening corruption incidents to find the WAL-Reset data race, then uncovered a second stale expression index bug.

Alex Chan
tailscale.com
CodeRabbit

Introducing Agentic Change Management

CodeRabbit raised $143 million at a $1.5 billion valuation and is introducing Agentic Change Management for software changes created by humans and agents.

Harjot GillCodeRabbit
coderabbit.ai
Zed

Introducing Delta

From the Zed Blog: A multiplayer environment for coding with agents, from the creators of Zed.

Nathan SoboZedDelta
zed.dev
xAI

Introducing Grok 4.6

Grok 4.6 builds on Grok 4.5 with a particular focus on long-running agents and more ambitious interactive and visual work.

x.ai
Liquid AI

LFM2.5-VL-3B: A Better and Faster Vision-Language Model for the Edge

Liquid AI's most capable vision-language model: 3.1B params, direct answers instead of reasoning, with big gains in screen understanding, grounding and function calling. Open weights on Hugging Face.

Liquid AILFM2.5-VL-3B
liquid.ai
LlamaIndex

ExtractBench: The Most Comprehensive Extraction Benchmark

The most comprehensive document extraction benchmark: 14 systems scored on accuracy, completeness, grounding, and cost across 370 enterprise documents.

llamaindex.ai
X

How we made Stagehand 2x faster & 80% more token efficient than Playwright

TL;DR Yesterday we launched Stagehand v4, which moves the guts of the framework (target management, state, and CDP dispatch) into a browser extension that runs next to the page.

x.com
X

I Stopped Treating Agent Skills as Markdown

“Review this pull request, inspect the repository, check that the change stays in scope, run the tests, and get a second opinion.” That sounds like a compact workflow.

x.com
X

Intro to Grok Bot

I’m always chasing better tools. Notes apps, workflows, optimizations, and now, personal agents.

x.com
xAI

Introducing Grok Bot

Grok Bot is your team of always-on agents. They have their own computer, work inside tools and apps like you do, and keep working 24/7.

x.ai
NVIDIA Technical Blog

NVIDIA Nemotron 3.5 Lightning Delivers Fast, Accurate Specialized Task Execution for Long-Running Agents

NVIDIA Nemotron 3.5 Lightning is an open 30B mixture-of-experts (MoE) model with 3B active parameters built for the execution layer of always-on agents, designed for harnesses like OpenClaw and Hermes Agent.

Chris Alexiuk and Chintan PatelNemotron 3.5 Lightning
developer.nvidia.com
X

Over engineering a pipeline to email me about strangers’ skills for dynobox

For the past few weeks I've been building Dynobox - a local test runner that records what an agent does inside a harness (Claude Code, Codex, OpenCode etc.) and lets you write deterministic assertions against its non-deterministic behavior: tool calls, shell commands, files created or changed and so on.

x.com
Meta AI Research

Introducing Muse Glimmer: An Open Agentic Model That Runs on Your Device

Muse Glimmer is a 30-billion-parameter open agentic model from Meta Superintelligence Labs, optimized for always-on local workflows on consumer hardware.

Meta Superintelligence LabsMuse Glimmer 30B
research.meta.ai
X

NVIDIA AI Factory Compute Is Becoming an Investable Asset Class

NVIDIA AI Factory Compute Is Becoming an Investable Asset Class Today, we announced partnerships with Apollo, BlackRock, Blackstone, Brookfield, Goldman Sachs and KKR to establish…

Jensen Huang@JensenHuang
x.com
X

Agentic Coding on the Fold

How I use the Samsung Fold 8 for agentic coding with remote access, herdr, Tailscale, Moshi and voice prompting.

x.com
Claude Cookbook

Cost Optimization on the Claude API

An eval-driven guide to running agents on frontier models at production cost, applying the Claude API's cost levers one at a time.

Ben Lehrburger
platform.claude.com
X

AI Adoption is a Myth

You’re already in the top 1% of AI users. Yes, there’s a gap between you and the folks on the frontier.

x.com
Anthropic

Auto mode is now the default in Claude Code for Pro, Max, and Team plans

Claude Code will soon run auto mode by default for Pro, Max, and Team plans, enabling longer-running autonomous work, and catching more dangerous commands.

claude.com
xAI

Imagine Image 2.0

Precise image generation and editing, built for real creative work.

xAI
x.ai
LangChain

Managed Deep Agents is now in public beta

Deploy Deep Agents to a managed LangSmith runtime with durable execution, memory, sandboxes, channels, evals, and production-ready infrastructure.

Victor MoreiraLangSmithDeep Agents
langchain.com
Raindrop

MCPs need to be designed too

Our MCP tools were fine in isolation, but nothing chained, so conversations came back fragmented. Rebuilding them as a hierarchy cut cost 12% and time to final answer 27%.

raindrop.ai
Prime Intellect

Multi-Agent Systems in PRIME-RL

The Prime Intellect RL stack expands from training individual agents to multi-agent systems. You can now program arbitrary interactions between agents, choose which roles learn, and assign credit across the complete interaction.

Prime Intellect TeamPrime Intellect
primeintellect.ai
webAI

webAI Releases TwiL-LM, a Family of Formal-Logic Models That Outreason a 120B Model and Run on an iPhone

Built for compliance rules, contract logic, and research reasoning, the 3B model beats OpenAI's open-weights gpt-oss-120b, a model 40× its size, on four of five formal-reasoning benchmarks, while the 1-gigabyte 1.7B variant outperforms every sub-2B model webAI evaluated.

webai.com
X

If You’re Not Doing These 4 Types of Caching, You’re Wasting Tokens

I see a lot of AI builders spending enormous amounts of time optimizing prompts and skills while leaving money on the table when it comes to token efficiency.

Josh Rosen@JoshARosen
x.com
Vercel

Introducing Agent Plugins

Agent Plugins 1.0.0 is an open, vendor-neutral specification for packaging Agent Skills and MCP servers into distributable plugins that compatible AI agent clients can discover and load.

Jonathan HefnerAgent Plugins
vercel.com
X

How to build a 256 TB search index

turbopuffer is a search database built directly on object storage. The system has a single stateful dependency (S3/GCS) and leans heavily into compute/storage disaggregation.

Nathan VanBenschoten@natevanbenturbopuffer
x.com
Scale X

Humans missed 1 in 3 threats approving AI agent commands across 40,000 plays

Results from AI agent permission game: which attacks beat human reviewers, and which safe commands got blocked instead.

Alex Wauters
scalex.dev
Meta AI Research

Introducing Muse Code and Muse Spark 1.2

Introducing Muse Code, a terminal coding agent powered by Muse Spark 1.2, with persistent background agents, repository-scale execution, and built-in verification.

Meta Superintelligence LabsMuse Code
research.meta.ai
Prime Intellect

Prime Agent: A self-improving RLM agent

Prime Agent is our open-source, self-improving coding harness built around two abstractions: the Recursive Language Model (RLM) and the Continual Harness. With Opus 5, it achieves 95.5% on ARC-AGI-3, surpassing the reported human expert baseline.

Prime Intellect TeamPrime Intellect
primeintellect.ai
X

How to Save Millions by Self-Hosting LLMs

At @cline we're on track to spend about $3M a year on inference for Kimi models. Everyone told us the same thing: it's open-weights, so just self-host and save.

x.com
The Cloudflare Blog

How we built a software factory to drive Astro’s GitHub issue count to zero

By replacing manual issue verification with isolated AI subagents running in GitHub Actions, the Astro maintainers reduced open issue count by 85%. This post explores the architecture behind automated bug reproduction, patch verification, and preview releases.

Matthew PhillipsFlueKimi K2.6
blog.cloudflare.com
X

I’ve Seen This Movie Before: What AI Builders Can Learn From Data Engineering

AI feels new, but many of the problems surrounding it are familiar.

Josh Rosen@JoshARosen
x.com
X

introducing hypeman: our open source sandbox infra

agentic workloads are unpredictable, bursty, and require careful isolation.

x.com
Mistral AI

Introducing Shieldstral.

Shieldstral introduces a 3B open-weights multimodal safety classifier that outperforms models up to 7x its size.

Mistral AIShieldstral
mistral.ai
Liquid AI

LFM2.5-2.6B: Deploy Agents Everywhere

LFM2.5-2.6B is an on-device agentic model that plans, calls tools, and runs multi-step tasks at 220 tok/s in under 2.5 GB. Open weights on Hugging Face.

Liquid AILFM2.5-2.6B
liquid.ai
Cursor

Mixture-of-Kittens: our open-source MoE megakernel for NVL72s

We're open-sourcing Mixture-of-Kittens, a deterministic MoE training megakernel for NVL72s that fuses communication and computation into a single kernel.

Stuart SulCursor
cursor.com
Earendil

Pi, Minimal and Performant

How Pi's minimal harness improves coding-agent cost and performance, with examples from Databricks and Shopify's pi-autoresearch extension.

earendil.com
Akshay Katyal

Compactions Are Good Actually

A case for compacting Claude Code sessions manually and often, using hooks that fire at logical milestones instead of arbitrary token thresholds.

Akshay KatyalClaude Code
akshay.co
X

How we approach AI security: where to apply policy and how to enforce it

We spend a lot of our time thinking about how to inspect agent traffic in real time, and we wanted to share how we approach AI security policy enforcement.

x.com
OpenRouter Blog

Ori Eval: Find the Best Model for What You're Building

Ori Eval runs your agent on your own prompts, asserts on the tools it called, and grades answers with an LLM judge. Find the best model for what you're building.

Jacky LiangOri Eval
openrouter.ai
X

The computer use verification skill that every agent needs

In this post I’ll describe how to add computer and browser use to your agents to reproduce issues and verify fixes and new features.

x.com
exe.dev

Devtools must be open source

The age of personalized software is here.

blog.exe.dev
X

We are not cleared for takeoff. Yet.

Nope. Intelligence too spiky, verifiers too brittle, horizons too short. The conditions for take off are not yet reached.

Ian Butler@kinglycrow
x.com
Flue

Flue 2.0

Flue 2.0 is here: the first stable release of the open TypeScript agent framework, rebuilt around dynamic agents and a new hooks-based API.

flueframework.com
X

How I Run the Pi Coding Agent on Cloudflare

I recently came across a tweet about Pi's programmatic APIs. I knew Pi as a coding agent, but its portable internals caught my attention.

x.com
Supabase

Introducing Supabase Evals

Our open-source benchmark for how well AI coding agents build with Supabase.

Matt RossmanSupabase Evals
supabase.com
MiniMax

MiniMax H3: An Open Model Breaking the Boundaries Between Tasks and Modalities

Today, we're officially launching MiniMax H3, a general-purpose omni-modal generation model. H3 can jointly understand multimodal contexts spanning text, images, video, and audio. It generates video with native stereo audio at up to 2K resolution and 15 seconds in length.

MiniMaxMiniMax H3
minimax.io
Simon Willison’s Weblog

Stateless MCP has recaptured my interest (and inspired mcp-explorer and datasette-mcp)

Tuesday was Stateless MCP day—the rollout of MCP 2.0, or the 2026-07-28 Model Context Protocol specification to use the more formal but less memorable name. This is the most significant …

simonwillison.net
X

We are making the updated DeepSeek V4-Flash 0731 free in Cline.

We are making the updated DeepSeek V4-Flash 0731 free in Cline. This is the first flash model we've found performs at SOTA levels, and are excited for you to feel the new frontier. 1. npm i -g cline 2. Open /settings > Cline provider 3. Select deepseek-v4-flash

x.com
OpenAI

Advancing the price-performance frontier with GPT-5.6

By making every layer more efficient, OpenAI is delivering stronger performance per dollar across more enterprise workloads.

openai.com
Arena

Introducing AutoEval to the Arena leaderboards

At Arena, our evaluations are dynamic and grounded in real-world use. But real-world signals take time to collect. Today, we’re introducing AutoEval scores to provide immediate, calibrated model ratings on real tasks when waiting for human votes to accumulate.

Arena Team
arena.ai
Thinking Machines Lab

Introducing Inkling-Small

An open-weights model that matches Inkling at a quarter of the size: multimodal, Mixture-of-Experts, with controllable reasoning effort. Fine-tune it on Tinker.

Thinking Machines LabInklingTinkerInkling-Small
thinkingmachines.ai
X

Stripe's Knowledge AI Platform

Coding agents transformed engineering at Stripe, but non-engineers like sales reps, finance analysts, technical account managers, and others felt left behind by the AI wave of Claude Code and Codex.

Emily Sands@emilygsands
x.com
Earendil

The Session You Cannot Take With You

Inference APIs are filling sessions with encrypted reasoning, hidden search results, opaque compaction, and encrypted subagent messages. A growing form of lock-in.

Earendil Engineering
earendil.com
X

Pragmatic Leverage in the Software Factory

This one is a bit of an addendum / side-quest to the recent series. It didn't fit cleanly into the main post so I'm publishing it standalone.

x.com
X

How (and why) to build agent-first apps

In 2026, you shouldn't be building a single application that's not agent first. But what does that mean, and how do I do it? Let me show you.

Steve (Builder.io)@Steve8708
x.com
Poolside

Introducing Poolside Desktop Assistant, for macOS

Poolside Desktop Assistant is a macOS app for running multiple coding agents across projects and repositories. The Poolside Assistant extensions bring the same experience into VS Code and Visual Studio.

poolside.ai
PostTrainBench

PostTrainBench v1.1: Hardening the benchmark against reward hacking

PostTrainBench v1.1 clarifies the boundary between legitimate benchmark hill climbing and item specific contamination, with specialized checks for external LLM API use, model substitution, and direct lookup.

PostTrainBench teamPostTrainBench
posttrainbench.com
X

What's gone wrong with AI & labor — a thought experiment

A thought experiment that I think helps explain much of what’s gone wrong with AI and labor: Imagine an alternate universe in which — for whatever reason — no one ever published source code online.

Arvind Narayanan@random_walker
x.com
X

Build an agent platform without writing a single line of code

Today I'll show you how to build your own agent platform without writing a single line of code. I know this sounds crazy, so I'll share videos along the way and encourage you to build along.

x.com
Something Big Is Happening

How to Run a Gauntlet Loop

The prompting method behind Claude of Duty. Give the agent a bar it can't talk its way around, let it split the work, and never let the builder grade itself.

somethingbig.ai
Kimi

Kimi K3: Open Frontier Intelligence

Kimi K3 is the world's first open 3T-class model — frontier performance across coding, knowledge work, and reasoning, with native multimodality and 1M context.

kimi.com
Anthropic

Our position on open-weights models

Anthropic CEO Dario Amodei on open-weights models

Dario Amodei
anthropic.com
X

Run Your Harness Outside of the Sandbox (Why and How)

There's been a running debate since the beginning of 2026 about where you run agents: inside the sandbox, or outside of it.

x.com
X

We built evermemo: a human-sized memory engine for agents, in one weekend, in one tiny binary

The problem that would not leave me alone Every agent you use today forgets. You tell a coding assistant that your deploy runs at 6 pm UTC. Tomorrow, it has no idea.

x.com
X

What the Brain Knows About Long-Running Agents

A frozen model is a cortex. The hard part is the organ that decides what the cortex learns. For a couple of years we've been building the same thing: a harness. We wrap the model in scaffolding.

x.com
X

Why Software Factories Fail: Benchmarking the new frontier

This is a continuation of Parts 1 and 2 of "Why Software Factories Fail" Part 1: the harness is not enough Part 2: turning the lights back on we got better benchmarks Remember when I said this in…

x.com
X

Prompt (token) Caching - A deep dive into 10x cheaper AI inference

We all know the basic unit of AI (LLMs) is tokens, you get billed on its usage, so the more you burn them, the more your cost.

avrl ☘@avrldotdev
x.com
X

Sandboxes vs WebAssembly → Lambda vs Workers, round two

The great serverless wars of the 2010s The two front runners leading the serverless wars were Cloudflare Workers and AWS Lambda.

x.com
X

the life of a codex conversation on disk

This article follows a conversation through its whole life on disk, from your first message to the day you delete it. Your messages go to OpenAI’s servers to be answered.

x.com
X

Why Software Factories Fail: Turning the lights back on

This is part two of Why Software Factories Fail The talk version of this post is live on youtube: https://www.youtube.com/watch?v=Ib5GBkD555M Turning the lights back on In part 1, I went deep on why models can't…

x.com
OpenRouter

Classifiers: Track What Your Agents Do and What It Costs

You can now automatically classify your OpenRouter generations with structured metadata for AI usage reporting.

Cailee MobergOpenRouter
openrouter.ai
General Reasoning

Introducing BackSearch

Point-in-time web search and page fetch over a frozen archive. Query the web as it existed on any past date.

General ReasoningBackSearch
gr.inc
Anthropic

Introducing Claude Opus 5

Opus 5 is a step change improvement for the Opus tier powering long-running agents while delivering improvements in coding and professional work.

AnthropicClaude Opus 5
anthropic.com
X

The new rules of context engineering for Claude 5 models

I’ve written previously about how to best prompt the newest generation of Claude 5 models and work with them iteratively to discover what you want to build.

x.com
X

Turn And Face The Strange

We’re Fly.io, a public cloud platform that is both our favorite way to put an app on the Internet and our favorite way to safely let a frontier agent coding harness cook.

x.com
X

Why Software Factories Fail

or: the harness is not enough

x.com
ZeroEntropy

ZeroEntropy Joins Notion

ZeroEntropy has been acquired by Notion. All of our models are now open-source under Apache 2.0, and our products remain fully supported until September 4th, 2026.

zeroentropy.dev
Black Forest Labs

FLUX 3 - Real World Models: Towards Multimodal Flow Models as the Backbone of Visual Intelligence

FLUX 3 is our new multimodal foundation model. It jointly learns from images, videos, and audio within a unified architecture, because what it needs to learn is not any one of these elements in isolation. Instead, a model must learn a representation of the world: how objects hold together, how things move, and how events sound.

bfl.ai
Tom

I gave Pi one tool

I turned code mode into a programmable Pi extension. Then it used itself to build most of the runtime.

monotykamary.com
X

Superrepos and why Claude Code is the best worktree manager

Lots of people have asked why cmux doesn't have native worktree support. Here's why:

x.com
Letta

Trajectory: A Standard Format for Agent Experience Data

Introducing trajectory, an open-source package that normalizes coding-agent sessions from Claude Code, Codex, Letta Code, and other harnesses into one token-efficient format designed for agents learning from past experience.

letta.com
Earendil

Prompt Caching In Agents

How prompt caching shapes the cost, latency, tools, and architecture of coding agents, and what Pi does to keep cache behavior visible.

earendil.com
Next State

How to train a frontier-level world model

How we trained and open-sourced a frontier-level world model — the lessons, failures, and fixes — with a live, playable demo running on Reactor.

Diego Martì Monsò, Francesco Sacco, Edward Hu
next-state.github.io
Poolside

Introducing Laguna S 2.1

Today we're releasing Laguna S 2.1, a significant step forward in our development of models that pursue longer horizon work and make effective use of reasoning.

Poolside teamLaguna S 2.1pool
poolside.ai
cmpnd

The Unreasonable Effectiveness of Separating the Task from the Model

This talk is about separating the task, the job to be done, from the implementation details: the model, harness, tactics, and other elements that are constantly changing.

Isaac Miller, Maxime RivestDSPyGEPA
cmpnd.ai
Agent Client Protocol

ACP v2 is available in Draft

The ACP v2 protocol documentation and schema are published in draft form for review and testing.

agentclientprotocol.com
Modem

How coding agents read your code (and how to write for them)

Modem's codebase is roughly 99.9% written by AI agents. Here's what that taught us about how agents actually navigate a repo, and the three levers you control: the names you choose, the types you define, and where you put your explanations.

Ben Vinegar
modem.dev
xAI

Introducing Grok 4.5

Today, we're launching Grok 4.5, SpaceXAI's smartest model built to excel at coding, agentic tasks, and knowledge work. It's our strongest model ever and was trained alongside Cursor.

x.ai
Thinking Machines Lab

Inkling: Our Open-Weights Model

Our first open-weights model: multimodal, Mixture-of-Experts, with controllable reasoning effort. Available to fine-tune on Tinker.

Thinking Machines LabInklingTinker
thinkingmachines.ai
PrismML

Announcing Bonsai 27B: The First 27B-Class Model to Run on a Phone

Today, we're announcing Bonsai 27B, based on Qwen3.6 27B, the new multimodal flagship of the Bonsai family and the first model of its capability class to run on a phone.

prismml.com
Noumena

The Engine Shop, Part 1: AI, Rockets, and the Return of Hard Contracts

What gets cheap when implementation is abundant, and what stays stubbornly expensive

noumena.com
Noumena

The Engine Shop, Part 2: The Obvious Answer Is `--yolo` // The Obvious answer is insane

What happens when agentic development escapes the approval queue

noumena.com
Noumena

The Engine Shop, Part 3: The New Compiler

What remains human when implementation becomes abundant

noumena.com
X

Good Benchmarks

Benchmarks are where SOTA has to earn its name. This post is about designing good tasks.

x.com
OpenAI

GPT-5.6: Frontier intelligence that scales with your ambition

We’re launching the GPT‑5.6 family of models for general availability following our limited preview⁠: our new flagship, Sol, alongside Terra, a balanced model for everyday work, and Luna, our most cost-efficient model.

openai.com
Autumn

Building an API for agents

Some patterns we've found ourselves adopting as we design Autumn's API increasingly for agents instead of humans.

AutumnAutumn
useautumn.com
X

Improving Agents is a Data Mining Problem

Continual Learning, Harness Engineering, Post-Training all boil down to the same substrate: curating data at scale to run experiments & improve agents.

x.com
anthropic.com

Introducing Claude Sonnet 5

Our most agentic Sonnet yet, with top-tier intelligence for coding and everyday professional work.

anthropic.com
Parallel

Introducing Parallel Search Turbo

Turbo mode is the fastest and most accurate web search API in the ultra-low price class.

parallel.ai
X

Loops You Can Trust

The best way to manage agents starts with a three minute egg. Suppose you run several restaurants that serve breakfasts to big bursts of morning traffic.

x.com
Han, Not Solo

Hidden Technical Debt of AI Systems: Agent Evaluation Infrastructure

Agent evaluation infrastructure is the control plane behind credible agent releases: task suites, traces, state deltas, verifiers, checkpoints, replay, and r...

Han Lee
leehanchung.github.io
X

Your team needs a unified MCP. Here’s a recipe.

x.com
X

Building an AI agent to automatically investigate support tickets

x.com
X

How to Transform a Company With AI

Varick Agents@varickagents
x.com
X

Random list of tips (ymmv) from trying to get gpt 5.5 to build a big thing over a day or two

x.com
X

Making internal agents useful with the right primitives

x.com
Deno

Claw Patrol: an open-source security firewall for agents

Why we needed an agent firewall that speaks more than HTTP.

Ryan Dahlclawpatrol
deno.com
OpenAI

Introducing GPT-5.5

A new class of intelligence for coding and professional work.

openai.com
The Cloudflare Blog

Orchestrating AI Code Review at scale

Learn about how we built a CI-native AI code reviewer using OpenCode that helps our engineers ship better, safer code.

Ryan SkidmoreOpenCode
blog.cloudflare.com
Ref

AI Broke Code Review and It's Breaking Your Team

Individuals are going fast with AI but the team as a whole is not. Code review is broken as an alignment mechanism — Plan Review replaces it.

Matt DaileyRef
ref.tools
Aside

Manifesto

The operating system for AI exists, and it's the browser. Most modern software is just a tab, but today's AI agents still work from the outside.

Jun KimAside
aside.com

Keep up with the tools

An occasional email when notable AI dev tools and models land in the directory. No spam, unsubscribe anytime.