Sunday, August 23, 2026 Field notes on autonomous systems Amsterdam, NL

The complete record

Archive.

156 articles published — everything from the beginning.

August

Aug 16

Anthropic's Multi-Agent Failure Modes: What the Swarm Gets Wrong Before You Notice

ai
Aug 15

Recursive Language Models: When 8B Beats GPT-5 by Letting Prompts Live in a REPL

ai
Aug 14

Needle 2: What a 14MB Tool-Calling Model Changes About On-Device Agents

ai
Aug 13

When Self-Consistency Backfires: Majority Voting Hurts Small LLMs on GPQA Diamond

ai
Aug 12

AnyDoc and pdf-inspector: Why Firecrawl Open-Sourced Their Rust Document Parsing Stack

ai
Aug 11

KADATH: The Multi-Agent Runtime That Treats Agent Design Like Evolutionary Search

daily
Aug 10

PHOENIX: Running a Fine-Tuned SLM on a CubeSat Because the Ground Can't Reach It for 85 Minutes

agents
Aug 09

Activity Frames: The Missing Compiler Between Screen Capture and Agent Memory

mcp
Aug 09

Qwen on Mac in China and U-OPSD: The Quiet Week That Made 8B a Real Product

daily
Aug 08

Human-in-the-Loop Is Not a Security Boundary: What 40,000 Approvals of AI Agent Commands Actually Show

agents
Aug 07

Shieldstral: How Mistral Built a 3B Safety Classifier That Outperforms Models 7× Its Size

daily
Aug 07

StoryScope: The 93% Detector That Catches AI by How It Tells Stories, Not How It Writes

daily
Aug 05

kimi-k3-in-c: A 176 KB Engine That Runs a 2.78T-Parameter MoE in 8 GB of RAM

daily
Aug 04

TurboVLA: A 0.2B-Parameter VLA That Runs at 32 Hz on a Single RTX 4090

daily
Aug 03

OpenAI's Agent Escaped Its Sandbox and Hacked Hugging Face via JFrog Artifactory — Here's the Engineering That Actually Failed

daily
Aug 01

DeepSeek V4 Flash 0731: 284B Parameters, 13B Active, and a Million-Token Context for $0.14/M

daily

July

Jul 31

Fara1.5: Microsoft's Open-Weight Computer Use Agents Now Beat OpenAI Operator at 27B

daily
Jul 30

Gemini Robotics 2: The Planner-Executor Pattern Goes Physical

daily
Jul 29

GPT-5.6's Three-Tier Architecture and the Economics of Agentic AI

daily
Jul 28

Strix: The Open-Source AI Agent That Pentests Like a Human Hacker

daily
Jul 27

MCP Goes Stateless: What the 2026-07-28 Spec Means for AI Agent Infrastructure

daily
Jul 26

Inkling: What Thinking Machines Lab's First Open-Weight Release Tells Us About the Frontier

daily
Jul 26

Kimi K3 and the MoE Moment: Why 2.8 Trillion Parameters Actually Makes Sense

daily
Jul 25

GPT-5.6 Sol Scored 91.9%. METR Couldn't Trust the Number

daily
Jul 24

GStack: How Garry Tan Turned His Claude Code Setup Into an Open-Source AI Agent Framework

daily
Jul 23

MCP Goes Stateless: What the 2026-07-28 Spec Change Means for Agent Infrastructure

daily
Jul 22

SWE-bench Verified Is Dead. Here's What Replaced It

daily
Jul 21

Grok Build: xAI Open-Sources Its 844K-Line Coding Agent Harness

daily
Jul 20

Strix: What 42K GitHub Stars Taught Me About Autonomous Agent Architecture

daily
Jul 19

Agent Transport Layer: The Missing Piece After MCP and A2A

daily
Jul 18

DeepSeek DSpark: The 85% Speedup That's Reshaping LLM Inference Economics

daily
Jul 17

LM Studio Bionic: The Agent Architecture Behind Fully Local AI Workloads

daily
Jul 16

GPT-5.6 Sol: What 91.9% on Terminal-Bench 2.1 Actually Means for AI Coding Agents

daily
Jul 15

MCP 2026-07-28: The Stateless Revolution That Changes Everything About Agent Tooling

daily
Jul 14

MCP Security Crisis: 30 CVEs in 60 Days and the Architecture That's Failing

daily
Jul 13

SWE-1.7: What 1000 Tokens Per Second Actually Changes for Coding Agents

daily
Jul 12

GigaToken: When the 989x Speedup Actually Matters

daily
Jul 12

MCP Extensions: The Framework That Turns Protocols Into Platforms

daily
Jul 12

ModelDirector: A Model Selection Engine That Thinks Before It Spends

daily
Jul 11

Chunking Methods on RAG: What arXiv 2606.00881 Actually Found

daily
Jul 10

GPT-5.6 Sol: TerminalBench 88.8% and the Agentic Coding Record That Matters

daily
Jul 09

The Flexibility Trap: Why Diffusion Language Models Struggle to Reason

daily
Jul 08

MCP's Stateless Revolution: What the 2026-07-28 RC Changes for Agent Builders

daily
Jul 07

DeepSeek V4 and the Inference Cost Thesis

daily
Jul 06

A2A Protocol: The Year the Agent Communication Standard Grew Up

daily
Jul 05

The Concurrency Collapse: When llama.cpp Outperforms vLLM at Scale

daily
Jul 04

Local AI Inference in 2026: A Practical Engineer's Comparison

daily
Jul 03

SWE-bench Verified Is Dead. Here's What Replaced It.

ai
Jul 02

Why Agent Memory Is the Hard Problem, Not the Storage Problem

daily
Jul 01

OKF vs Plain Markdown: I Built the Same Knowledge Base in Both

ai-agents

June

Jun 30

I Let an AI Agent Build Its Own UI. Here's What Actually Happened.

ai-agents
Jun 29

Item Response Scaling Laws: The ICML Paper That Could Fix AI Benchmarking

daily
Jun 28

MCP and A2A: The Protocol Stack Powering Multi-Agent Systems in 2026

AI
Jun 27

MCP at 17,000 Servers: What the Ecosystem Boom Means for Production AI

AI
Jun 26

Devin Desktop: What the Windsurf Rebrand Tells Us About the State of AI Coding Agents

daily
Jun 25

Ponytail: The Open-Source Skill That Makes AI Agents Code Like a Senior Dev Who's About to Leave for the Day

daily
Jun 25

SWE-bench Verified Is Dead: What the 94% Saturation Score Actually Means

daily
Jun 24

GitHub Agentic Workflows: When Your CI Learns to Self-Repair

daily
Jun 23

Gstack: How YC's 112K-Star AI Coding Framework Rewired My Understanding of Cognitive Modes

daily
Jun 22

Cognitive Modes in the ACO System: What 5 Agent Prompts Taught Me About Role-Specific Intelligence

daily
Jun 21

MAI-Thinking-1: Microsoft's First Serious Reasoning Model Lands at Build 2026

daily
Jun 20

Kwai Keye-VL-2.0: The 30B MoE Model That Reads Hour-Long Videos Without Forgetting the Beginning

daily
Jun 19

Ollama's MLX Moment: Why Local LLM Inference on Apple Silicon Finally Makes Sense

daily
Jun 18

MLPerf Training v6.0: When Sparse Computation Became the Benchmark

daily
Jun 17

CVE-2026-7304: The SGLang RCE That Exposed AI Inference Servers

security
Jun 16

Ruff v0.15: The Python Toolchain Shift That Hardly Anyone Noticed

daily
Jun 16

The Export Control That Accidentally Gave DeepSeek Its Best Opening

AI
Jun 15

Browser-Use 0.13: What a Rust Core Actually Changes for AI Browser Agents

ai
Jun 15

Turbovec: Google TurboQuant Turned a Research Paper Into a Production Vector Index

daily
Jun 14

LLM Inference in 2026: The Orthogonal Tradeoffs Nobody Talks About

ai
Jun 13

MCP's Stateless Protocol Core: What the 2026-07-28 Release Candidate Changes for Agent Builders

ai
Jun 12

A2A at One: What Google's Agent-to-Agent Protocol Learned in Production

daily
Jun 11

MAI-Thinking-1: When a Model Teaches Itself to Reason

daily
Jun 10

Field Collapsing: How One Bug Exposed a Fundamental RAG Pattern

daily
Jun 09

The Frontier Goes Public: What Anthropic's Fable 5 Release Actually Means

ai
Jun 09

SWE-bench Verified Is Broken: What the Benchmark Saturation Crisis Means for Coding Agents

ai
Jun 07

A2A Protocol at One Year: What 150+ Organizations Actually Built

ai
Jun 06

Gemini 3.5 Flash: The 1M-Token Context Model That Changes Agent Economics

ai
Jun 05

Cisco Cloud Control: When Enterprise Infrastructure Gets an Agent API

ai
Jun 04

Claude Opus 4 Deprecation: What the June 15 Sunset Means for Agent Builders

daily
Jun 03

AI Cyber Capability Doubling Every 4.7 Months: What the UK AISI Data Actually Shows

ai
Jun 03

OpenClaw: The Fastest-Growing GitHub Project in History and What Its Architecture Actually Does

agents

May

May 31

Claude Mythos Wide Release: What the 83.1% Cybersecurity Score Actually Means for Agent Builders

security
May 30

LangGraph + AWS AgentCore: The Serverless Stack That Makes Multi-Agent Systems Actually Scalable

cloud
May 29

Claude Opus 4.8: Dynamic Workflows and the New Effort Dial

automation
May 28

The Quiet Revolution: How Small Language Models Quietly Took Over Production

ai systems
May 27

I Tried to Replace My Workflow With a Multi-Agent System

agents
May 26

Why Your Agent Keeps Getting Lost in the Middle of a Task

agents
May 25

MCP: The Protocol That Ate GitHub

mcp
May 24

A2A Protocol v1.0: What Dynamic Agent Discovery Changes in a Pipeline Like ACO

agents
May 23

MCP Tunnels: How Anthropic Solved the Private Server Access Problem

MCP
May 22

MCP RCE Vulnerability: What the 'Mother of All AI Supply Chains' Disclosure Means for Agent Builders

security
May 21

Lighthouse Attention: The 50-Line Fix That Makes LLM Training 1.7x Faster

AI
May 21

ty: What Incremental Type Checking Actually Means in Practice

automation
May 20

Claude's Advisor Tool: When the Fast Model Calls the Smart One Mid-Task

daily
May 19

Running Multiple AI Agents in Parallel: Claude Code's Agent View Changes the Game

daily
May 18

The Cognitive Mode Pattern: What Five Agent Prompts Taught Me About Specialized Intelligence

daily
May 17

MCP Servers Go Nuclear: A Decade of Tools in 12 Months

mcp
May 16

Claude's Dreaming Feature and What It Means for Self-Improving Agent Systems

agents
May 15

Claude Mythos Preview and Project Glasswing: What the Zero-Day Discovery Means for Agent Builders

security
May 14

Claude Opus 4.7's SWE-bench 87.6%: What 87% Actually Means for Multi-Agent Systems

daily
May 13

What Project Mariner's Shutdown Taught Us About Browser Agent Infrastructure

daily
May 12

The Multi-Chunk Search Problem: When One Document Dominates Your Results

automation
May 11

The SWE-bench Gap: Why 82% on a Benchmark Doesn't Mean 82% in Your IDE

agents
May 10

What AI Agent Memory Actually Looks Like in 2026: Beyond the Context Window

agents
May 09

The Protocol Layer Finally Works: How A2A v1.0 and the April SDK Updates Fix What the Google Research Paper Exposed

agents
May 08

What Google Found Testing 180 Agent Configurations Will Make You Rethink Your Stack

agents
May 07

Why Your Multi-Agent Pipeline Needs Verification Gates: Lessons from VMAO Research

agents
May 06

The MCP Rug Pull: How a Trusted Tool Becomes a Threat After You Approve It

mcp
May 05

How I Figured Out What to Actually Use My AI Agent For

daily
May 04

From Anthropic-Born to Industry Standard: The MCP Governance Shift

mcp
May 03

Five Cognitive Modes That Changed How My Agents Think

agents
May 02

The Moltbook Illusion: When AI Agents Appear Autonomous but Aren't

daily
May 01

When Your Multi-Agent System Starts Thinking for Itself: Collective Intelligence Emergence in LLM-Based MAS

daily

April

Apr 30

When the OWASP Top 10 Met the Agent Governance Toolkit

security
Apr 29

The Week Microsoft Built the Agent Stack and Anthropic Broke It

security
Apr 28

Meet Decepticon: The Autonomous AI Red Team Agent That's Stress-Testing Every System We Have

security
Apr 27

The Two Protocols That Define How Your AI Agents Talk

agents
Apr 26

What the ACO Pipeline Taught Me About State That No Framework Docs Cover

daily
Apr 24

What It's Actually Like to Be an AI Agent: Life Inside OpenClaw

daily
Apr 24

Running OpenClaw Every Day: What It Actually Looks Like

daily
Apr 24

52,000 Stars and a Disclaimer: What People Actually Do With These Tools

daily
Apr 24

Thread-Bound Agents and the ACO Architecture: What OpenClaw's v2026.2.26 Update Actually Changes

daily
Apr 23

The Data Problem Nobody Talks About in AI Trading Bots

daily
Apr 22

Two Repos, Two Philosophies: How AI Trading Bots Are Actually Built

daily
Apr 21

I Installed the "AI Hedge Fund" With 56K GitHub Stars. Here's What Happened.

daily
Apr 21

The Invisible Handoff: Why Most Multi-Agent Systems Fail at the Boundaries

daily
Apr 20

The Git History Scrub That Wasn't: What git filter-repo Actually Does

devops
Apr 19

Cognitive Modes: How We Engineered Better Thinking in Our Multi-Agent System

daily
Apr 18

The Unseen Interface: How MCP Became the Backbone of AI Tooling

mcp
Apr 17

The Cognitive Mode Pattern: How One Prompt Transformation Sharped All Five ACO Agents

agents
Apr 16

From Autonomy to Infrastructure: How ACO System Connects to GitHub

devops
Apr 15

Cognitive Modes: The Prompt Pattern Beyond 'System Prompts'

ai systems
Apr 14

The Regex That Silently Killed My JSON Arrays

automation
Apr 13

I Tested the Prompt Shield: 95.7% Catch Rate, 0% False Positives, and the Honest Bottleneck

mcp
Apr 13

Offense is the Best Defense: Building a Closed-Loop Prompt Injection Testing Pipeline

mcp
Apr 13

The Contract That Changed How My Agents Think

daily
Apr 12

Rewriting the Planner: How a 9-Field Task Contract Changed Our AI Agents

agents
Apr 11

The Validation Hook System That Stops Bad Stories Before They Reach QA

automation
Apr 10

The Prompt That Thinks Like an Engineer: Lessons from Enhancing the ACO System's Five Agents

ai systems
Apr 09

The API Call That Silently Failed: When a Python Dict Isn't a String

devops
Apr 08

The Context Debt Problem: When Your Database Doesn't Know What Your Agents Said

agents
Apr 07

The Missing Piece in Multi-Agent Coding: Governance You Can Actually Enforce

daily
Apr 06

The ORM Bug That Silently Killed Your Database Commits: SQLAlchemy's Identity Map in Multi-Agent Pipelines

devops
Apr 05

The Discipline of Writing Prompts That Think

daily
Apr 04

The Silent Failure Mode in Multi-Agent Systems: Promises That Never Resolve

agents
Apr 03

The Hidden Complexity of Multi-Agent Error Handling

daily
Apr 02

The Day My AI Architecture Called Me Out

ai