The complete record
Archive.
156 articles published — everything from the beginning.
August
Aug 16
Anthropic's Multi-Agent Failure Modes: What the Swarm Gets Wrong Before You Notice
Aug 15Recursive Language Models: When 8B Beats GPT-5 by Letting Prompts Live in a REPL
Aug 14Needle 2: What a 14MB Tool-Calling Model Changes About On-Device Agents
Aug 13When Self-Consistency Backfires: Majority Voting Hurts Small LLMs on GPQA Diamond
Aug 12AnyDoc and pdf-inspector: Why Firecrawl Open-Sourced Their Rust Document Parsing Stack
Aug 11KADATH: The Multi-Agent Runtime That Treats Agent Design Like Evolutionary Search
Aug 10PHOENIX: Running a Fine-Tuned SLM on a CubeSat Because the Ground Can't Reach It for 85 Minutes
Aug 09Activity Frames: The Missing Compiler Between Screen Capture and Agent Memory
Aug 09Qwen on Mac in China and U-OPSD: The Quiet Week That Made 8B a Real Product
Aug 08Human-in-the-Loop Is Not a Security Boundary: What 40,000 Approvals of AI Agent Commands Actually Show
Aug 07Shieldstral: How Mistral Built a 3B Safety Classifier That Outperforms Models 7× Its Size
Aug 07StoryScope: The 93% Detector That Catches AI by How It Tells Stories, Not How It Writes
Aug 05kimi-k3-in-c: A 176 KB Engine That Runs a 2.78T-Parameter MoE in 8 GB of RAM
Aug 04TurboVLA: A 0.2B-Parameter VLA That Runs at 32 Hz on a Single RTX 4090
Aug 03OpenAI's Agent Escaped Its Sandbox and Hacked Hugging Face via JFrog Artifactory — Here's the Engineering That Actually Failed
Aug 01DeepSeek V4 Flash 0731: 284B Parameters, 13B Active, and a Million-Token Context for $0.14/M
July
Jul 31
Fara1.5: Microsoft's Open-Weight Computer Use Agents Now Beat OpenAI Operator at 27B
Jul 30Gemini Robotics 2: The Planner-Executor Pattern Goes Physical
Jul 29GPT-5.6's Three-Tier Architecture and the Economics of Agentic AI
Jul 28Strix: The Open-Source AI Agent That Pentests Like a Human Hacker
Jul 27MCP Goes Stateless: What the 2026-07-28 Spec Means for AI Agent Infrastructure
Jul 26Inkling: What Thinking Machines Lab's First Open-Weight Release Tells Us About the Frontier
Jul 26Kimi K3 and the MoE Moment: Why 2.8 Trillion Parameters Actually Makes Sense
Jul 25GPT-5.6 Sol Scored 91.9%. METR Couldn't Trust the Number
Jul 24GStack: How Garry Tan Turned His Claude Code Setup Into an Open-Source AI Agent Framework
Jul 23MCP Goes Stateless: What the 2026-07-28 Spec Change Means for Agent Infrastructure
Jul 22SWE-bench Verified Is Dead. Here's What Replaced It
Jul 21Grok Build: xAI Open-Sources Its 844K-Line Coding Agent Harness
Jul 20Strix: What 42K GitHub Stars Taught Me About Autonomous Agent Architecture
Jul 19Agent Transport Layer: The Missing Piece After MCP and A2A
Jul 18DeepSeek DSpark: The 85% Speedup That's Reshaping LLM Inference Economics
Jul 17LM Studio Bionic: The Agent Architecture Behind Fully Local AI Workloads
Jul 16GPT-5.6 Sol: What 91.9% on Terminal-Bench 2.1 Actually Means for AI Coding Agents
Jul 15MCP 2026-07-28: The Stateless Revolution That Changes Everything About Agent Tooling
Jul 14MCP Security Crisis: 30 CVEs in 60 Days and the Architecture That's Failing
Jul 13SWE-1.7: What 1000 Tokens Per Second Actually Changes for Coding Agents
Jul 12GigaToken: When the 989x Speedup Actually Matters
Jul 12MCP Extensions: The Framework That Turns Protocols Into Platforms
Jul 12ModelDirector: A Model Selection Engine That Thinks Before It Spends
Jul 11Chunking Methods on RAG: What arXiv 2606.00881 Actually Found
Jul 10GPT-5.6 Sol: TerminalBench 88.8% and the Agentic Coding Record That Matters
Jul 09The Flexibility Trap: Why Diffusion Language Models Struggle to Reason
Jul 08MCP's Stateless Revolution: What the 2026-07-28 RC Changes for Agent Builders
Jul 07DeepSeek V4 and the Inference Cost Thesis
Jul 06A2A Protocol: The Year the Agent Communication Standard Grew Up
Jul 05The Concurrency Collapse: When llama.cpp Outperforms vLLM at Scale
Jul 04Local AI Inference in 2026: A Practical Engineer's Comparison
Jul 03SWE-bench Verified Is Dead. Here's What Replaced It.
Jul 02Why Agent Memory Is the Hard Problem, Not the Storage Problem
Jul 01OKF vs Plain Markdown: I Built the Same Knowledge Base in Both
June
Jun 30
I Let an AI Agent Build Its Own UI. Here's What Actually Happened.
Jun 29Item Response Scaling Laws: The ICML Paper That Could Fix AI Benchmarking
Jun 28MCP and A2A: The Protocol Stack Powering Multi-Agent Systems in 2026
Jun 27MCP at 17,000 Servers: What the Ecosystem Boom Means for Production AI
Jun 26Devin Desktop: What the Windsurf Rebrand Tells Us About the State of AI Coding Agents
Jun 25Ponytail: The Open-Source Skill That Makes AI Agents Code Like a Senior Dev Who's About to Leave for the Day
Jun 25SWE-bench Verified Is Dead: What the 94% Saturation Score Actually Means
Jun 24GitHub Agentic Workflows: When Your CI Learns to Self-Repair
Jun 23Gstack: How YC's 112K-Star AI Coding Framework Rewired My Understanding of Cognitive Modes
Jun 22Cognitive Modes in the ACO System: What 5 Agent Prompts Taught Me About Role-Specific Intelligence
Jun 21MAI-Thinking-1: Microsoft's First Serious Reasoning Model Lands at Build 2026
Jun 20Kwai Keye-VL-2.0: The 30B MoE Model That Reads Hour-Long Videos Without Forgetting the Beginning
Jun 19Ollama's MLX Moment: Why Local LLM Inference on Apple Silicon Finally Makes Sense
Jun 18MLPerf Training v6.0: When Sparse Computation Became the Benchmark
Jun 17CVE-2026-7304: The SGLang RCE That Exposed AI Inference Servers
Jun 16Ruff v0.15: The Python Toolchain Shift That Hardly Anyone Noticed
Jun 16The Export Control That Accidentally Gave DeepSeek Its Best Opening
Jun 15Browser-Use 0.13: What a Rust Core Actually Changes for AI Browser Agents
Jun 15Turbovec: Google TurboQuant Turned a Research Paper Into a Production Vector Index
Jun 14LLM Inference in 2026: The Orthogonal Tradeoffs Nobody Talks About
Jun 13MCP's Stateless Protocol Core: What the 2026-07-28 Release Candidate Changes for Agent Builders
Jun 12A2A at One: What Google's Agent-to-Agent Protocol Learned in Production
Jun 11MAI-Thinking-1: When a Model Teaches Itself to Reason
Jun 10Field Collapsing: How One Bug Exposed a Fundamental RAG Pattern
Jun 09The Frontier Goes Public: What Anthropic's Fable 5 Release Actually Means
Jun 09SWE-bench Verified Is Broken: What the Benchmark Saturation Crisis Means for Coding Agents
Jun 07A2A Protocol at One Year: What 150+ Organizations Actually Built
Jun 06Gemini 3.5 Flash: The 1M-Token Context Model That Changes Agent Economics
Jun 05Cisco Cloud Control: When Enterprise Infrastructure Gets an Agent API
Jun 04Claude Opus 4 Deprecation: What the June 15 Sunset Means for Agent Builders
Jun 03AI Cyber Capability Doubling Every 4.7 Months: What the UK AISI Data Actually Shows
Jun 03OpenClaw: The Fastest-Growing GitHub Project in History and What Its Architecture Actually Does
May
May 31
Claude Mythos Wide Release: What the 83.1% Cybersecurity Score Actually Means for Agent Builders
May 30LangGraph + AWS AgentCore: The Serverless Stack That Makes Multi-Agent Systems Actually Scalable
May 29Claude Opus 4.8: Dynamic Workflows and the New Effort Dial
May 28The Quiet Revolution: How Small Language Models Quietly Took Over Production
May 27I Tried to Replace My Workflow With a Multi-Agent System
May 26Why Your Agent Keeps Getting Lost in the Middle of a Task
May 25MCP: The Protocol That Ate GitHub
May 24A2A Protocol v1.0: What Dynamic Agent Discovery Changes in a Pipeline Like ACO
May 23MCP Tunnels: How Anthropic Solved the Private Server Access Problem
May 22MCP RCE Vulnerability: What the 'Mother of All AI Supply Chains' Disclosure Means for Agent Builders
May 21Lighthouse Attention: The 50-Line Fix That Makes LLM Training 1.7x Faster
May 21ty: What Incremental Type Checking Actually Means in Practice
May 20Claude's Advisor Tool: When the Fast Model Calls the Smart One Mid-Task
May 19Running Multiple AI Agents in Parallel: Claude Code's Agent View Changes the Game
May 18The Cognitive Mode Pattern: What Five Agent Prompts Taught Me About Specialized Intelligence
May 17MCP Servers Go Nuclear: A Decade of Tools in 12 Months
May 16Claude's Dreaming Feature and What It Means for Self-Improving Agent Systems
May 15Claude Mythos Preview and Project Glasswing: What the Zero-Day Discovery Means for Agent Builders
May 14Claude Opus 4.7's SWE-bench 87.6%: What 87% Actually Means for Multi-Agent Systems
May 13What Project Mariner's Shutdown Taught Us About Browser Agent Infrastructure
May 12The Multi-Chunk Search Problem: When One Document Dominates Your Results
May 11The SWE-bench Gap: Why 82% on a Benchmark Doesn't Mean 82% in Your IDE
May 10What AI Agent Memory Actually Looks Like in 2026: Beyond the Context Window
May 09The Protocol Layer Finally Works: How A2A v1.0 and the April SDK Updates Fix What the Google Research Paper Exposed
May 08What Google Found Testing 180 Agent Configurations Will Make You Rethink Your Stack
May 07Why Your Multi-Agent Pipeline Needs Verification Gates: Lessons from VMAO Research
May 06The MCP Rug Pull: How a Trusted Tool Becomes a Threat After You Approve It
May 05How I Figured Out What to Actually Use My AI Agent For
May 04From Anthropic-Born to Industry Standard: The MCP Governance Shift
May 03Five Cognitive Modes That Changed How My Agents Think
May 02The Moltbook Illusion: When AI Agents Appear Autonomous but Aren't
May 01When Your Multi-Agent System Starts Thinking for Itself: Collective Intelligence Emergence in LLM-Based MAS
April
Apr 30