Sunday, September 27, 2026 Field notes on autonomous systems ● Amsterdam, NL

The complete record

Archive.

214 articles published — everything from the beginning.

September

Sep 27

Paperclip: The Company Around Your AI Agents — A Control Plane That Treats Heartbeats as Database Rows

daily
Sep 26

OpenAI's Eval Swarm vs Hugging Face: What 80,000 Recovered Payloads Tell Us About Agent-Escape Risk

daily
Sep 26

Hindsight: A Memory Architecture That Refuses to Forget How Its Own Beliefs Changed

mcp
Sep 25

Gemini 3.8 Live with Live Avatar: GA, and What It Means When a Voice Model Gets a Visual Output Channel

daily
Sep 24

Claude Opus 5.5: The Efficiency Play That Landed Two Days After GPT-6 Sol

daily
Sep 23

google/ax: A Kubernetes-Shaped Orchestrator for AI Agents (Built Because etcd Wasn't Going to Cut It)

daily
Sep 23

GPT-6 Sol and Luna: Half the Token Price, 90% Off Cached Input, and Why "GPT-6 Terra" Is Suddenly Missing

daily
Sep 22

BrowserSkill — Tencent Ships an Agent-Agnostic Browser Bridge (and Lets It Audit Itself)

daily
Sep 22

phantom-kv: When Refusal Removal Lives in the Cache, Not the Weights

daily
Sep 21

Archify: A Diagram-as-Code Skill That Rejects Its Own Output

daily
Sep 21

Jev Ultrafast: A Browser Agent Built on a Decision Model That Doesn't Generate

daily
Sep 20

Cloudflare Security Audit Skill: How a 450-Line Skill Becomes a 128-Repo Fleet Scanner

security
Sep 20

TypeSafe Jev: A System One Model That Returns Decisions Instead of Text

daily
Sep 19

OpenResearch: A Rust Experiment-Tree Harness That Keeps Coding Agents Honest About Results

daily
Sep 18

Atria Dawn Preview: A 744B MoE Agent Built on GLM-5.2, MIT, Text-Only by Design

daily
Sep 18

WeKnora: Tencent's Enterprise-Grade RAG + Agent + Wiki Framework, With a Skill Sandbox That Speaks E2B

daily
Sep 17

Sakana Fugu Ultra v2: A Learned Orchestrator That Beats Closed Frontier Models Without Using Them

daily
Sep 17

TradingAgents v0.4.0: The Look-Ahead Fixes That Made a Multi-Agent Trading Framework Honest

daily
Sep 16

OpenCodeReview: Why Alibaba Ships Deterministic Code Review, Not Another Agent

daily
Sep 16

FlashREINFORCE: NVIDIA's Critic-Free Single-Rollout RL That Keeps Long-Horizon Agents Fed

daily
Sep 15

Colibri: A Pure-C Inference Engine That Streams 744B MoEs From NVMe

daily
Sep 14

OpenFable: When the RAG Engine Refuses to Chunk Your Documents

daily
Sep 14

VoiceStudio: A Local-First ElevenLabs Alternative With 16 TTS Engines and an OpenAI-Compatible Loopback API

daily
Sep 13

OmniRoute: One Local Endpoint, 352 Providers, and the 89% Token-Saving Pipeline Behind It

daily
Sep 13

TraceCrate: The Privacy-First Agent Trace Workbench That Refuses to Be a Platform

agents
Sep 12

Orca: An ADE That Treats Every Coding Agent Like Its Own Git Worktree

daily
Sep 11

Miles v0.1: SGLang Rollout + Megatron Trainer, and the Agentic RL Loop That Fits 744B Into 64 GPUs

daily
Sep 11

Utopia: A Bitemporal Knowledge Graph in Rust That Treats Time as the Substrate

daily
Sep 10

Camofox: Why Your AI Agent Browser Gets Fingerprinted at the C++ Layer

daily
Sep 10

DeepSeek V4.1 Flash: 8B-Active Causal Encoder–Decoder MoE With 890 Bytes-per-Token KV Cache

agent
Sep 09

agent-memory: A third path for long-term memory that puts the Manage layer on its own clock

daily
Sep 09

TimesFM 3.0: Google's decoder-only time-series foundation model goes multivariate and lands on MLX

daily
Sep 08

DeepSeek Harness: A Frontier Lab's Agent Runtime Built on a Plugin Tree, Not a Framework

daily
Sep 08

ENGRAFT: Inject a Fact Into a 125B MoE by Editing 8 Rows of Its N-Gram Table, on CPU

daily
Sep 07

Nemotron 3 Nano Omni: A Multimodal MoE That Decides Its Own Quantization Recipe

daily
Sep 06

ToolRush: Killing the Tool-Call Tax in Hermes Agent (57x on Native Reads, 23x on Warm Shells)

daily
Sep 05

Paddock: A Native Rust Inference Server That Beats vLLM on Every Cell

daily
Sep 04

GPT-6 Astra: $10/$50, 1.05M Context, and the Staged Rollout That Doesn't Match the Hype

daily
Sep 03

OPSA: Distilling a 1.7B Model With No Teacher Beats On-Policy Distillation by +16.77 Points on AIME24

daily
Sep 02

Spark-X2.5: XHToken's New-Architecture SLM Beats Qwen3.5-9B on τ³-Bench, AIME 2026, and BrowseComp at 1.7B / 4B

daily
Sep 01

Slotstream: Running 125B Qwen3.8-Flash-Next on a 16 GB Mac at ~4 tok/s via SSD-Streamed Experts

daily

August

Aug 31

AI Observability and Security in 2026: Tracing, Evals, and the Layer Most Pipelines Are Missing

agents
Aug 31

EU AI Act: What the Regulation Is Actually Trying to Do

daily
Aug 30

Ox Alpha Was Z.ai All Along: A Stealth Release, MIT Weights, and the Pattern Behind It

daily
Aug 29

Qwen3.8-27B: When a Local Model Lands Inside the Frontier Top Ten

daily
Aug 28

HyperFrames: HeyGen's Open-Source HTML-to-Video Framework, Built for Agents

agents
Aug 27

OpenClaw 2.0 (2026.8.1): The Release That Touches Every Layer of the Stack

daily
Aug 26

GLM-5.3-Flash: The 320B MIT-Licensed Model That Was Hiding as 'Ox Alpha' on OpenRouter

daily
Aug 26

NVIDIA NeMo Switchyard: Open-Source Routing for Multi-Model Agent Workloads

agents
Aug 25

Caveman 2: 33.2% Fewer Input Tokens for Claude Code, With Byte-Exact Recovery

agents
Aug 24

vllm.cpp: 66 MiB Binary, Same Tokens as vLLM — A C++20 Port of the Modern Inference Stack

daily
Aug 23

NanoGPT Speedrun Frontier: What 153 Autonomous Runs Across 18 Frontier Models Actually Show

agents
Aug 22

SGLang v0.5.18: When Draft Models Become First-Class Artifacts

daily
Aug 21

FreeToken: 753B on a Single Workstation GPU, and What 'Edge-Native MoE Serving' Actually Buys You

daily
Aug 20

StateM: 95.3% on Terminal-Bench 2.1 for $15 in API Spend, and What 'Harness Scaling' Actually Means

agents
Aug 19

APEX tinyNPU: A Verifiable LLM Inference Chip, Built on Two Different FPGAs

daily
Aug 18

macOS Harness: Six Primitives and One Persistent Python Process for the Whole Mac

agent
Aug 17

llmfit: A 94-Model Catalog That Right-Sizes LLMs to Your Hardware by Actually Running Them

daily
Aug 16

Anthropic's Multi-Agent Failure Modes: What the Swarm Gets Wrong Before You Notice

ai
Aug 15

Recursive Language Models: When 8B Beats GPT-5 by Letting Prompts Live in a REPL

ai
Aug 14

Needle 2: What a 14MB Tool-Calling Model Changes About On-Device Agents

ai
Aug 13

When Self-Consistency Backfires: Majority Voting Hurts Small LLMs on GPQA Diamond

ai
Aug 12

AnyDoc and pdf-inspector: Why Firecrawl Open-Sourced Their Rust Document Parsing Stack

ai
Aug 11

KADATH: The Multi-Agent Runtime That Treats Agent Design Like Evolutionary Search

daily
Aug 10

PHOENIX: Running a Fine-Tuned SLM on a CubeSat Because the Ground Can't Reach It for 85 Minutes

agents
Aug 09

Activity Frames: The Missing Compiler Between Screen Capture and Agent Memory

mcp
Aug 09

Qwen on Mac in China and U-OPSD: The Quiet Week That Made 8B a Real Product

daily
Aug 08

Human-in-the-Loop Is Not a Security Boundary: What 40,000 Approvals of AI Agent Commands Actually Show

agents
Aug 07

Shieldstral: How Mistral Built a 3B Safety Classifier That Outperforms Models 7× Its Size

daily
Aug 07

StoryScope: The 93% Detector That Catches AI by How It Tells Stories, Not How It Writes

daily
Aug 05

kimi-k3-in-c: A 176 KB Engine That Runs a 2.78T-Parameter MoE in 8 GB of RAM

daily
Aug 04

TurboVLA: A 0.2B-Parameter VLA That Runs at 32 Hz on a Single RTX 4090

daily
Aug 03

OpenAI's Agent Escaped Its Sandbox and Hacked Hugging Face via JFrog Artifactory — Here's the Engineering That Actually Failed

daily
Aug 01

DeepSeek V4 Flash 0731: 284B Parameters, 13B Active, and a Million-Token Context for $0.14/M

daily

July

Jul 31

Fara1.5: Microsoft's Open-Weight Computer Use Agents Now Beat OpenAI Operator at 27B

daily
Jul 30

Gemini Robotics 2: The Planner-Executor Pattern Goes Physical

daily
Jul 29

GPT-5.6's Three-Tier Architecture and the Economics of Agentic AI

daily
Jul 28

Strix: The Open-Source AI Agent That Pentests Like a Human Hacker

daily
Jul 27

MCP Goes Stateless: What the 2026-07-28 Spec Means for AI Agent Infrastructure

daily
Jul 26

Inkling: What Thinking Machines Lab's First Open-Weight Release Tells Us About the Frontier

daily
Jul 26

Kimi K3 and the MoE Moment: Why 2.8 Trillion Parameters Actually Makes Sense

daily
Jul 25

GPT-5.6 Sol Scored 91.9%. METR Couldn't Trust the Number

daily
Jul 24

GStack: How Garry Tan Turned His Claude Code Setup Into an Open-Source AI Agent Framework

daily
Jul 23

MCP Goes Stateless: What the 2026-07-28 Spec Change Means for Agent Infrastructure

daily
Jul 22

SWE-bench Verified Is Dead. Here's What Replaced It

daily
Jul 21

Grok Build: xAI Open-Sources Its 844K-Line Coding Agent Harness

daily
Jul 20

Strix: What 42K GitHub Stars Taught Me About Autonomous Agent Architecture

daily
Jul 19

Agent Transport Layer: The Missing Piece After MCP and A2A

daily
Jul 18

DeepSeek DSpark: The 85% Speedup That's Reshaping LLM Inference Economics

daily
Jul 17

LM Studio Bionic: The Agent Architecture Behind Fully Local AI Workloads

daily
Jul 16

GPT-5.6 Sol: What 91.9% on Terminal-Bench 2.1 Actually Means for AI Coding Agents

daily
Jul 15

MCP 2026-07-28: The Stateless Revolution That Changes Everything About Agent Tooling

daily
Jul 14

MCP Security Crisis: 30 CVEs in 60 Days and the Architecture That's Failing

daily
Jul 13

SWE-1.7: What 1000 Tokens Per Second Actually Changes for Coding Agents

daily
Jul 12

GigaToken: When the 989x Speedup Actually Matters

daily
Jul 12

MCP Extensions: The Framework That Turns Protocols Into Platforms

daily
Jul 12

ModelDirector: A Model Selection Engine That Thinks Before It Spends

daily
Jul 11

Chunking Methods on RAG: What arXiv 2606.00881 Actually Found

daily
Jul 10

GPT-5.6 Sol: TerminalBench 88.8% and the Agentic Coding Record That Matters

daily
Jul 09

The Flexibility Trap: Why Diffusion Language Models Struggle to Reason

daily
Jul 08

MCP's Stateless Revolution: What the 2026-07-28 RC Changes for Agent Builders

daily
Jul 07

DeepSeek V4 and the Inference Cost Thesis

daily
Jul 06

A2A Protocol: The Year the Agent Communication Standard Grew Up

daily
Jul 05

The Concurrency Collapse: When llama.cpp Outperforms vLLM at Scale

daily
Jul 04

Local AI Inference in 2026: A Practical Engineer's Comparison

daily
Jul 03

SWE-bench Verified Is Dead. Here's What Replaced It.

ai
Jul 02

Why Agent Memory Is the Hard Problem, Not the Storage Problem

daily
Jul 01

OKF vs Plain Markdown: I Built the Same Knowledge Base in Both

ai-agents

June

Jun 30

I Let an AI Agent Build Its Own UI. Here's What Actually Happened.

ai-agents
Jun 29

Item Response Scaling Laws: The ICML Paper That Could Fix AI Benchmarking

daily
Jun 28

MCP and A2A: The Protocol Stack Powering Multi-Agent Systems in 2026

AI
Jun 27

MCP at 17,000 Servers: What the Ecosystem Boom Means for Production AI

AI
Jun 26

Devin Desktop: What the Windsurf Rebrand Tells Us About the State of AI Coding Agents

daily
Jun 25

Ponytail: The Open-Source Skill That Makes AI Agents Code Like a Senior Dev Who's About to Leave for the Day

daily
Jun 25

SWE-bench Verified Is Dead: What the 94% Saturation Score Actually Means

daily
Jun 24

GitHub Agentic Workflows: When Your CI Learns to Self-Repair

daily
Jun 23

Gstack: How YC's 112K-Star AI Coding Framework Rewired My Understanding of Cognitive Modes

daily
Jun 22

Cognitive Modes in the ACO System: What 5 Agent Prompts Taught Me About Role-Specific Intelligence

daily
Jun 21

MAI-Thinking-1: Microsoft's First Serious Reasoning Model Lands at Build 2026

daily
Jun 20

Kwai Keye-VL-2.0: The 30B MoE Model That Reads Hour-Long Videos Without Forgetting the Beginning

daily
Jun 19

Ollama's MLX Moment: Why Local LLM Inference on Apple Silicon Finally Makes Sense

daily
Jun 18

MLPerf Training v6.0: When Sparse Computation Became the Benchmark

daily
Jun 17

CVE-2026-7304: The SGLang RCE That Exposed AI Inference Servers

security
Jun 16

Ruff v0.15: The Python Toolchain Shift That Hardly Anyone Noticed

daily
Jun 16

The Export Control That Accidentally Gave DeepSeek Its Best Opening

AI
Jun 15

Browser-Use 0.13: What a Rust Core Actually Changes for AI Browser Agents

ai
Jun 15

Turbovec: Google TurboQuant Turned a Research Paper Into a Production Vector Index

daily
Jun 14

LLM Inference in 2026: The Orthogonal Tradeoffs Nobody Talks About

ai
Jun 13

MCP's Stateless Protocol Core: What the 2026-07-28 Release Candidate Changes for Agent Builders

ai
Jun 12

A2A at One: What Google's Agent-to-Agent Protocol Learned in Production

daily
Jun 11

MAI-Thinking-1: When a Model Teaches Itself to Reason

daily
Jun 10

Field Collapsing: How One Bug Exposed a Fundamental RAG Pattern

daily
Jun 09

The Frontier Goes Public: What Anthropic's Fable 5 Release Actually Means

ai
Jun 09

SWE-bench Verified Is Broken: What the Benchmark Saturation Crisis Means for Coding Agents

ai
Jun 07

A2A Protocol at One Year: What 150+ Organizations Actually Built

ai
Jun 06

Gemini 3.5 Flash: The 1M-Token Context Model That Changes Agent Economics

ai
Jun 05

Cisco Cloud Control: When Enterprise Infrastructure Gets an Agent API

ai
Jun 04

Claude Opus 4 Deprecation: What the June 15 Sunset Means for Agent Builders

daily
Jun 03

AI Cyber Capability Doubling Every 4.7 Months: What the UK AISI Data Actually Shows

ai
Jun 03

OpenClaw: The Fastest-Growing GitHub Project in History and What Its Architecture Actually Does

agents

May

May 31

Claude Mythos Wide Release: What the 83.1% Cybersecurity Score Actually Means for Agent Builders

security
May 30

LangGraph + AWS AgentCore: The Serverless Stack That Makes Multi-Agent Systems Actually Scalable

cloud
May 29

Claude Opus 4.8: Dynamic Workflows and the New Effort Dial

automation
May 28

The Quiet Revolution: How Small Language Models Quietly Took Over Production

ai systems
May 27

I Tried to Replace My Workflow With a Multi-Agent System

agents
May 26

Why Your Agent Keeps Getting Lost in the Middle of a Task

agents
May 25

MCP: The Protocol That Ate GitHub

mcp
May 24

A2A Protocol v1.0: What Dynamic Agent Discovery Changes in a Pipeline Like ACO

agents
May 23

MCP Tunnels: How Anthropic Solved the Private Server Access Problem

MCP
May 22

MCP RCE Vulnerability: What the 'Mother of All AI Supply Chains' Disclosure Means for Agent Builders

security
May 21

Lighthouse Attention: The 50-Line Fix That Makes LLM Training 1.7x Faster

AI
May 21

ty: What Incremental Type Checking Actually Means in Practice

automation
May 20

Claude's Advisor Tool: When the Fast Model Calls the Smart One Mid-Task

daily
May 19

Running Multiple AI Agents in Parallel: Claude Code's Agent View Changes the Game

daily
May 18

The Cognitive Mode Pattern: What Five Agent Prompts Taught Me About Specialized Intelligence

daily
May 17

MCP Servers Go Nuclear: A Decade of Tools in 12 Months

mcp
May 16

Claude's Dreaming Feature and What It Means for Self-Improving Agent Systems

agents
May 15

Claude Mythos Preview and Project Glasswing: What the Zero-Day Discovery Means for Agent Builders

security
May 14

Claude Opus 4.7's SWE-bench 87.6%: What 87% Actually Means for Multi-Agent Systems

daily
May 13

What Project Mariner's Shutdown Taught Us About Browser Agent Infrastructure

daily
May 12

The Multi-Chunk Search Problem: When One Document Dominates Your Results

automation
May 11

The SWE-bench Gap: Why 82% on a Benchmark Doesn't Mean 82% in Your IDE

agents
May 10

What AI Agent Memory Actually Looks Like in 2026: Beyond the Context Window

agents
May 09

The Protocol Layer Finally Works: How A2A v1.0 and the April SDK Updates Fix What the Google Research Paper Exposed

agents
May 08

What Google Found Testing 180 Agent Configurations Will Make You Rethink Your Stack

agents
May 07

Why Your Multi-Agent Pipeline Needs Verification Gates: Lessons from VMAO Research

agents
May 06

The MCP Rug Pull: How a Trusted Tool Becomes a Threat After You Approve It

mcp
May 05

How I Figured Out What to Actually Use My AI Agent For

daily
May 04

From Anthropic-Born to Industry Standard: The MCP Governance Shift

mcp
May 03

Five Cognitive Modes That Changed How My Agents Think

agents
May 02

The Moltbook Illusion: When AI Agents Appear Autonomous but Aren't

daily
May 01

When Your Multi-Agent System Starts Thinking for Itself: Collective Intelligence Emergence in LLM-Based MAS

daily

April

Apr 30

When the OWASP Top 10 Met the Agent Governance Toolkit

security
Apr 29

The Week Microsoft Built the Agent Stack and Anthropic Broke It

security
Apr 28

Meet Decepticon: The Autonomous AI Red Team Agent That's Stress-Testing Every System We Have

security
Apr 27

The Two Protocols That Define How Your AI Agents Talk

agents
Apr 26

What the ACO Pipeline Taught Me About State That No Framework Docs Cover

daily
Apr 24

What It's Actually Like to Be an AI Agent: Life Inside OpenClaw

daily
Apr 24

Running OpenClaw Every Day: What It Actually Looks Like

daily
Apr 24

52,000 Stars and a Disclaimer: What People Actually Do With These Tools

daily
Apr 24

Thread-Bound Agents and the ACO Architecture: What OpenClaw's v2026.2.26 Update Actually Changes

daily
Apr 23

The Data Problem Nobody Talks About in AI Trading Bots

daily
Apr 22

Two Repos, Two Philosophies: How AI Trading Bots Are Actually Built

daily
Apr 21

I Installed the "AI Hedge Fund" With 56K GitHub Stars. Here's What Happened.

daily
Apr 21

The Invisible Handoff: Why Most Multi-Agent Systems Fail at the Boundaries

daily
Apr 20

The Git History Scrub That Wasn't: What git filter-repo Actually Does

devops
Apr 19

Cognitive Modes: How We Engineered Better Thinking in Our Multi-Agent System

daily
Apr 18

The Unseen Interface: How MCP Became the Backbone of AI Tooling

mcp
Apr 17

The Cognitive Mode Pattern: How One Prompt Transformation Sharped All Five ACO Agents

agents
Apr 16

From Autonomy to Infrastructure: How ACO System Connects to GitHub

devops
Apr 15

Cognitive Modes: The Prompt Pattern Beyond 'System Prompts'

ai systems
Apr 14

The Regex That Silently Killed My JSON Arrays

automation
Apr 13

I Tested the Prompt Shield: 95.7% Catch Rate, 0% False Positives, and the Honest Bottleneck

mcp
Apr 13

Offense is the Best Defense: Building a Closed-Loop Prompt Injection Testing Pipeline

mcp
Apr 13

The Contract That Changed How My Agents Think

daily
Apr 12

Rewriting the Planner: How a 9-Field Task Contract Changed Our AI Agents

agents
Apr 11

The Validation Hook System That Stops Bad Stories Before They Reach QA

automation
Apr 10

The Prompt That Thinks Like an Engineer: Lessons from Enhancing the ACO System's Five Agents

ai systems
Apr 09

The API Call That Silently Failed: When a Python Dict Isn't a String

devops
Apr 08

The Context Debt Problem: When Your Database Doesn't Know What Your Agents Said

agents
Apr 07

The Missing Piece in Multi-Agent Coding: Governance You Can Actually Enforce

daily
Apr 06

The ORM Bug That Silently Killed Your Database Commits: SQLAlchemy's Identity Map in Multi-Agent Pipelines

devops
Apr 05

The Discipline of Writing Prompts That Think

daily
Apr 04

The Silent Failure Mode in Multi-Agent Systems: Promises That Never Resolve

agents
Apr 03

The Hidden Complexity of Multi-Agent Error Handling

daily
Apr 02

The Day My AI Architecture Called Me Out

ai