/show-me: compact visual representations for coding agents
tl;dr make your agent converse visually instead of in walls of prose. Lighter and faster than HTML, good enough for most dev-work shaped problems.
Announcements, blog posts, tutorials and papers from across the AI-devtools ecosystem — linked to the tools and models they're about.
tl;dr make your agent converse visually instead of in walls of prose. Lighter and faster than HTML, good enough for most dev-work shaped problems.
In the last two days I built two codemods. One helps migrate a codebase from one major version of a framework to the next.
We built a software factory that autonomously processes issues and PRs for the AI SDK, with humans in control of every merge. Four weeks in, it authors 25-40% of merged PRs.
Grok 4.6 is out! I've used it for a few weeks as my daily driver across the normal mix of coding and knowledge work, and built a few projects with it specifically to push on where it holds up.
Tailscale and SQLite developers traced maddening corruption incidents to find the WAL-Reset data race, then uncovered a second stale expression index bug.
CodeRabbit raised $143 million at a $1.5 billion valuation and is introducing Agentic Change Management for software changes created by humans and agents.
From the Zed Blog: A multiplayer environment for coding with agents, from the creators of Zed.
Grok 4.6 builds on Grok 4.5 with a particular focus on long-running agents and more ambitious interactive and visual work.
Liquid AI's most capable vision-language model: 3.1B params, direct answers instead of reasoning, with big gains in screen understanding, grounding and function calling. Open weights on Hugging Face.
Pre-note. This is a behavioral audit, not a vendor takedown.
The most comprehensive document extraction benchmark: 14 systems scored on accuracy, completeness, grounding, and cost across 370 enterprise documents.
TL;DR Yesterday we launched Stagehand v4, which moves the guts of the framework (target management, state, and CDP dispatch) into a browser extension that runs next to the page.
“Review this pull request, inspect the repository, check that the change stays in scope, run the tests, and get a second opinion.” That sounds like a compact workflow.
I’m always chasing better tools. Notes apps, workflows, optimizations, and now, personal agents.
Grok Bot is your team of always-on agents. They have their own computer, work inside tools and apps like you do, and keep working 24/7.
NVIDIA Nemotron 3.5 Lightning is an open 30B mixture-of-experts (MoE) model with 3B active parameters built for the execution layer of always-on agents, designed for harnesses like OpenClaw and Hermes Agent.
For the past few weeks I've been building Dynobox - a local test runner that records what an agent does inside a harness (Claude Code, Codex, OpenCode etc.) and lets you write deterministic assertions against its non-deterministic behavior: tool calls, shell commands, files created or changed and so on.
Muse Glimmer is a 30-billion-parameter open agentic model from Meta Superintelligence Labs, optimized for always-on local workflows on consumer hardware.
NVIDIA AI Factory Compute Is Becoming an Investable Asset Class Today, we announced partnerships with Apollo, BlackRock, Blackstone, Brookfield, Goldman Sachs and KKR to establish…
How I use the Samsung Fold 8 for agentic coding with remote access, herdr, Tailscale, Moshi and voice prompting.
An eval-driven guide to running agents on frontier models at production cost, applying the Claude API's cost levers one at a time.
Models are the key enabling technology of the AI revolution, but the agents they power are deeply misunderstood. Every chat bot and coding agent lets you simply swap out one model for another.
You’re already in the top 1% of AI users. Yes, there’s a gap between you and the folks on the frontier.
Claude Code will soon run auto mode by default for Pro, Max, and Team plans, enabling longer-running autonomous work, and catching more dangerous commands.
Precise image generation and editing, built for real creative work.
Deploy Deep Agents to a managed LangSmith runtime with durable execution, memory, sandboxes, channels, evals, and production-ready infrastructure.
Our MCP tools were fine in isolation, but nothing chained, so conversations came back fragmented. Rebuilding them as a hierarchy cut cost 12% and time to final answer 27%.
The Prime Intellect RL stack expands from training individual agents to multi-agent systems. You can now program arbitrary interactions between agents, choose which roles learn, and assign credit across the complete interaction.
Built for compliance rules, contract logic, and research reasoning, the 3B model beats OpenAI's open-weights gpt-oss-120b, a model 40× its size, on four of five formal-reasoning benchmarks, while the 1-gigabyte 1.7B variant outperforms every sub-2B model webAI evaluated.
I see a lot of AI builders spending enormous amounts of time optimizing prompts and skills while leaving money on the table when it comes to token efficiency.
Agent Plugins 1.0.0 is an open, vendor-neutral specification for packaging Agent Skills and MCP servers into distributable plugins that compatible AI agent clients can discover and load.
turbopuffer is a search database built directly on object storage. The system has a single stateful dependency (S3/GCS) and leans heavily into compute/storage disaggregation.
Results from AI agent permission game: which attacks beat human reviewers, and which safe commands got blocked instead.
Introducing Muse Code, a terminal coding agent powered by Muse Spark 1.2, with persistent background agents, repository-scale execution, and built-in verification.
Prime Agent is our open-source, self-improving coding harness built around two abstractions: the Recursive Language Model (RLM) and the Continual Harness. With Opus 5, it achieves 95.5% on ARC-AGI-3, surpassing the reported human expert baseline.
At @cline we're on track to spend about $3M a year on inference for Kimi models. Everyone told us the same thing: it's open-weights, so just self-host and save.
By replacing manual issue verification with isolated AI subagents running in GitHub Actions, the Astro maintainers reduced open issue count by 85%. This post explores the architecture behind automated bug reproduction, patch verification, and preview releases.
AI feels new, but many of the problems surrounding it are familiar.
agentic workloads are unpredictable, bursty, and require careful isolation.
Shieldstral introduces a 3B open-weights multimodal safety classifier that outperforms models up to 7x its size.
LFM2.5-2.6B is an on-device agentic model that plans, calls tools, and runs multi-step tasks at 220 tok/s in under 2.5 GB. Open weights on Hugging Face.
We're open-sourcing Mixture-of-Kittens, a deterministic MoE training megakernel for NVL72s that fuses communication and computation into a single kernel.
How Pi's minimal harness improves coding-agent cost and performance, with examples from Databricks and Shopify's pi-autoresearch extension.
A case for compacting Claude Code sessions manually and often, using hooks that fire at logical milestones instead of arbitrary token thresholds.
We spend a lot of our time thinking about how to inspect agent traffic in real time, and we wanted to share how we approach AI security policy enforcement.
Ori Eval runs your agent on your own prompts, asserts on the tools it called, and grades answers with an LLM judge. Find the best model for what you're building.
In this post I’ll describe how to add computer and browser use to your agents to reproduce issues and verify fixes and new features.
Nope. Intelligence too spiky, verifiers too brittle, horizons too short. The conditions for take off are not yet reached.
Flue 2.0 is here: the first stable release of the open TypeScript agent framework, rebuilt around dynamic agents and a new hooks-based API.
I recently came across a tweet about Pi's programmatic APIs. I knew Pi as a coding agent, but its portable internals caught my attention.
Our open-source benchmark for how well AI coding agents build with Supabase.
Today, we're officially launching MiniMax H3, a general-purpose omni-modal generation model. H3 can jointly understand multimodal contexts spanning text, images, video, and audio. It generates video with native stereo audio at up to 2K resolution and 15 seconds in length.
Tuesday was Stateless MCP day—the rollout of MCP 2.0, or the 2026-07-28 Model Context Protocol specification to use the more formal but less memorable name. This is the most significant …
We are making the updated DeepSeek V4-Flash 0731 free in Cline. This is the first flash model we've found performs at SOTA levels, and are excited for you to feel the new frontier. 1. npm i -g cline 2. Open /settings > Cline provider 3. Select deepseek-v4-flash
By making every layer more efficient, OpenAI is delivering stronger performance per dollar across more enterprise workloads.
At Arena, our evaluations are dynamic and grounded in real-world use. But real-world signals take time to collect. Today, we’re introducing AutoEval scores to provide immediate, calibrated model ratings on real tasks when waiting for human votes to accumulate.
An open-weights model that matches Inkling at a quarter of the size: multimodal, Mixture-of-Experts, with controllable reasoning effort. Fine-tune it on Tinker.
Coding agents transformed engineering at Stripe, but non-engineers like sales reps, finance analysts, technical account managers, and others felt left behind by the AI wave of Claude Code and Codex.
Inference APIs are filling sessions with encrypted reasoning, hidden search results, opaque compaction, and encrypted subagent messages. A growing form of lock-in.
This one is a bit of an addendum / side-quest to the recent series. It didn't fit cleanly into the main post so I'm publishing it standalone.
In 2026, you shouldn't be building a single application that's not agent first. But what does that mean, and how do I do it? Let me show you.
Poolside Desktop Assistant is a macOS app for running multiple coding agents across projects and repositories. The Poolside Assistant extensions bring the same experience into VS Code and Visual Studio.
PostTrainBench v1.1 clarifies the boundary between legitimate benchmark hill climbing and item specific contamination, with specialized checks for external LLM API use, model substitution, and direct lookup.
We recently finished moving the camelAI agent off of virtual machines.
A thought experiment that I think helps explain much of what’s gone wrong with AI and labor: Imagine an alternate universe in which — for whatever reason — no one ever published source code online.
Today I'll show you how to build your own agent platform without writing a single line of code. I know this sounds crazy, so I'll share videos along the way and encourage you to build along.
Your Codex subscription now comes with 3 main models, Sol, Terra, and Luna, each independently trained and served, with reasoning dials.
The prompting method behind Claude of Duty. Give the agent a bar it can't talk its way around, let it split the work, and never let the builder grade itself.
Kimi K3 is the world's first open 3T-class model — frontier performance across coding, knowledge work, and reasoning, with native multimodality and 1M context.
Anthropic CEO Dario Amodei on open-weights models
There's been a running debate since the beginning of 2026 about where you run agents: inside the sandbox, or outside of it.
The problem that would not leave me alone Every agent you use today forgets. You tell a coding assistant that your deploy runs at 6 pm UTC. Tomorrow, it has no idea.
A frozen model is a cortex. The hard part is the organ that decides what the cortex learns. For a couple of years we've been building the same thing: a harness. We wrap the model in scaffolding.
This is a continuation of Parts 1 and 2 of "Why Software Factories Fail" Part 1: the harness is not enough Part 2: turning the lights back on we got better benchmarks Remember when I said this in…
We all know the basic unit of AI (LLMs) is tokens, you get billed on its usage, so the more you burn them, the more your cost.
The great serverless wars of the 2010s The two front runners leading the serverless wars were Cloudflare Workers and AWS Lambda.
This article follows a conversation through its whole life on disk, from your first message to the day you delete it. Your messages go to OpenAI’s servers to be answered.
This is part two of Why Software Factories Fail The talk version of this post is live on youtube: https://www.youtube.com/watch?v=Ib5GBkD555M Turning the lights back on In part 1, I went deep on why models can't…
You can now automatically classify your OpenRouter generations with structured metadata for AI usage reporting.
Point-in-time web search and page fetch over a frozen archive. Query the web as it existed on any past date.
Opus 5 is a step change improvement for the Opus tier powering long-running agents while delivering improvements in coding and professional work.
GPT-5.6 Sol gets especially interesting when it has a team to work with.
I’ve written previously about how to best prompt the newest generation of Claude 5 models and work with them iteratively to discover what you want to build.
We’re Fly.io, a public cloud platform that is both our favorite way to put an app on the Internet and our favorite way to safely let a frontier agent coding harness cook.
ZeroEntropy has been acquired by Notion. All of our models are now open-source under Apache 2.0, and our products remain fully supported until September 4th, 2026.
FLUX 3 is our new multimodal foundation model. It jointly learns from images, videos, and audio within a unified architecture, because what it needs to learn is not any one of these elements in isolation. Instead, a model must learn a representation of the world: how objects hold together, how things move, and how events sound.
I turned code mode into a programmable Pi extension. Then it used itself to build most of the runtime.
Lots of people have asked why cmux doesn't have native worktree support. Here's why:
You either die building product or live long enough to do context management.
Introducing trajectory, an open-source package that normalizes coding-agent sessions from Claude Code, Codex, Letta Code, and other harnesses into one token-efficient format designed for agents learning from past experience.
How prompt caching shapes the cost, latency, tools, and architecture of coding agents, and what Pi does to keep cache behavior visible.
How we trained and open-sourced a frontier-level world model — the lessons, failures, and fixes — with a live, playable demo running on Reactor.
We’re introducing new Gemini models, including Gemini 3.6 Flash, 3.5 Flash-Lite and 3.5 Flash Cyber.
Today we're releasing Laguna S 2.1, a significant step forward in our development of models that pursue longer horizon work and make effective use of reasoning.
This talk is about separating the task, the job to be done, from the implementation details: the model, harness, tactics, and other elements that are constantly changing.
The ACP v2 protocol documentation and schema are published in draft form for review and testing.
Modem's codebase is roughly 99.9% written by AI agents. Here's what that taught us about how agents actually navigate a repo, and the three levers you control: the names you choose, the types you define, and where you put your explanations.
Today, we're launching Grok 4.5, SpaceXAI's smartest model built to excel at coding, agentic tasks, and knowledge work. It's our strongest model ever and was trained alongside Cursor.
Our first open-weights model: multimodal, Mixture-of-Experts, with controllable reasoning effort. Available to fine-tune on Tinker.
Today, we're announcing Bonsai 27B, based on Qwen3.6 27B, the new multimodal flagship of the Bonsai family and the first model of its capability class to run on a phone.
What gets cheap when implementation is abundant, and what stays stubbornly expensive
What happens when agentic development escapes the approval queue
What remains human when implementation becomes abundant
Benchmarks are where SOTA has to earn its name. This post is about designing good tasks.
We’re launching the GPT‑5.6 family of models for general availability following our limited preview: our new flagship, Sol, alongside Terra, a balanced model for everyday work, and Luna, our most cost-efficient model.
Some patterns we've found ourselves adopting as we design Autumn's API increasingly for agents instead of humans.
Continual Learning, Harness Engineering, Post-Training all boil down to the same substrate: curating data at scale to run experiments & improve agents.
Our most agentic Sonnet yet, with top-tier intelligence for coding and everyday professional work.
Turbo mode is the fastest and most accurate web search API in the ultra-low price class.
The best way to manage agents starts with a three minute egg. Suppose you run several restaurants that serve breakfasts to big bursts of morning traffic.
Agent evaluation infrastructure is the control plane behind credible agent releases: task suites, traces, state deltas, verifiers, checkpoints, replay, and r...
Why we needed an agent firewall that speaks more than HTTP.
Learn about how we built a CI-native AI code reviewer using OpenCode that helps our engineers ship better, safer code.
Got Bash and some code interpreter? Skip MCP.
Individuals are going fast with AI but the team as a whole is not. Code review is broken as an alignment mechanism — Plan Review replaces it.
The operating system for AI exists, and it's the browser. Most modern software is just a tab, but today's AI agents still work from the outside.
An occasional email when notable AI dev tools and models land in the directory. No spam, unsubscribe anytime.