google/ax: A Kubernetes-Shaped Orchestrator for AI Agents (Built Because etcd Wasn't Going to Cut It) — aniketkarneai.com | aniketkarneai.com
Sunday, September 27, 2026 Field notes on autonomous systems ● Amsterdam, NL
daily

google/ax: A Kubernetes-Shaped Orchestrator for AI Agents (Built Because etcd Wasn't Going to Cut It)

Google's open-source AX ships as a Kubernetes-shaped control plane for agentic workloads — Task, Workspace, Gateway, and Model primitives stored in Redis Streams instead of CRDs because etcd collapses under millions of short-lived tasks. v0.3.0 landed on Sep 20 with a generic-purpose reorg and pinned workflow actions.

The README’s first paragraph says the quiet part out loud: “We are still actively refining our core concepts, protocols, and specifications. We will likely to introduce major breaking changes prior to a stable release.” That warning block sits above a WARNING callout, which is what Google puts on infrastructure it intends to keep moving. The repo is at v0.3.0, the latest tag landed on 2026-09-20 (three days ago at the time of writing), and the most recent commit on main is the one that restructured AX into “a general-purpose orchestration layer for agentic tasks” — a deliberate reorg away from whatever the project was first scoped as.

I want to read the v0.3.0 design carefully because it is the rare 2026 infrastructure release from a major lab that publishes the control-plane internals — Redis Streams instead of etcd, an atespace concept that is not a Kubernetes namespace, a custom runner contract, and an explicit “you can replace the runner with your own binary” API. Most agent-framework releases stop at “apply YAML, watch a thing happen.” AX ships the design rationale along with the manifests.

What AX actually is, stripped of Kubernetes phrasing

The pitch is: declare an agentic task with workspaces and gateway specifications, and AX sandboxes it, wires up its workspace, fences its network, and helps running it at scale. The README borrows kubectl vocabulary deliberately — ax apply, ax get, ax describe, ax watch, ax delete, plus a few agent-specific verbs (ax ssh, ax suspend, ax resume). The CLI is shaped the way Kubernetes operators expect because the maintainers want kubectl people to feel at home without learning a new mental model.

But the architectural shape is not Kubernetes. There are four primitives, and they are smaller than CRDs:

  • Task — the smallest unit of isolated execution. Declares the container image, command, compute requests and limits, env vars, a Gateway reference, and one or more Workspace references. Lifecycle: Running, Suspended, Failed, Terminating, with conditions WorkspaceReady, GatewayReady, Ready.
  • Workspace — populates the filesystem and tool landscape before the agent starts. Clones Git repos at the right revisions, sets up MCP servers and skill registries, and exposes a goal field that gets handed to a bootstrap agent on first boot (the default image uses Antigravity with GEMINI_API_KEY from the atespace’s Model config).
  • Gateway — the network boundary. Listeners the task exposes, plus an egress allowlist of hosts and ports the sandbox may reach. This is where you fence an agent to “your LLM provider and your Git host” and nothing else.
  • Model — not a model. A named model configuration: which provider, which model identifier, provider-specific generation parameters, and a Kubernetes secretKeyRef holding the API key. The configuration lives in one place instead of in every agent’s environment.

Everything lives in an atespace. The default atespace is default, but you can scope per-team or per-tenant. The name is a Googleism and worth flagging — it is not a Kubernetes namespace, although it occupies roughly the same conceptual role. ax -a my-atespace get tasks scopes a query; the atespace appears in the ate-target-actor: <atespace>/<task> header that the network router resolves.

The README is candid that the unit is deliberately small: “An agent is not one process that runs to completion; over its lifetime it plans, delegates, retries, and fans work out. AX does not try to model that shape. It gives you one primitive that is cheap to create, isolate, suspend, and throw away, and lets the agent compose as many of them as its work demands.” That sentence is doing real work — it is a refusal of the “agent as long-lived runtime” framing that competes in the same space as the open-source Agent Substrate runtime that AX runs on top of.

Why Redis Streams and not Kubernetes CRDs

This is the part of DESIGN.md I keep coming back to. The opening paragraph is unusually honest:

Storing millions of short-lived tasks as Kubernetes CRDs pushes etcd past its comfort zone (single-digit GB storage limits, write-rate bottlenecks, control plane degradation). AX keeps its state in Redis and uses Redis Streams as the work queue between the API server and a horizontally scaled pool of controllers.

That is a real engineering claim, not a benchmark. The control-plane diagram in DESIGN.md shows the topology plainly: ax apply → ax-server (stateless gRPC API on port 8080) → store and publish event → Redis (Task Hashes + Event Streams + PubSub) → XREADGROUP on Streams → ax-controller (horizontally scaled workers) → gRPC to Agent Substrate. The state plane is Redis; the work queue is Redis Streams with consumer groups; the controllers do reconciliation. None of this passes through kube-apiserver, none of this writes a CRD, and none of this eats an etcd lease per task.

The implication for someone used to K8s operators is that watching a task is not the same as watching a CRD. The watch path is server-streaming gRPC on WatchTask, not a kube informer’s cache. Polling patterns that work against CRDs (label selectors, field selectors, kubectl get -w) don’t transfer. ax watch task task123 is the operator’s primary tool, and it streams status.phase plus condition transitions as they happen.

The four binaries in the design map onto that topology:

BinaryRole
axDeveloper CLI. Applies manifests, inspects and watches resources, tunnels to the cluster.
ax-serverStateless gRPC API on port 8080. Validates manifests, persists to Redis, publishes events.
ax-controllerReconciliation workers. Consume the Redis stream, provision atespaces and actors on Agent Substrate, apply egress policy, and drive tasks toward desired state. Scale by adding replicas.
ax-task-runnerEntrypoint inside every task container. Bootstraps the workspace, serves metadata, and runs the agent command.

The horizontal scaling story is also clean: you add ax-controller replicas to consume the Redis Stream faster, you don’t have to thread etcd leases. The control plane stays stateless, and the bottleneck becomes Redis throughput rather than etcd’s serialized write path. For a workload class where every agent break-down produces a tree of subtasks, that distinction matters — kubectl run and forget about it is a fine story for a pod that lives an hour, but agents fan out work and the fan-out compounds.

The runner contract is the part most agent frameworks skip

The single best document in the repo is docs/runner.md. It defines what a runner must do, in detail, and then says: “AX ships a default runner, ax-task-runner, baked into the default task image. You do not have to use it.” That sentence is doing the same kind of deliberate framing that the task-as-small-unit quote is doing. The control plane tells the agent what to do; the runner tells the container how to do it. Either side is replaceable.

The contract is concrete:

  • Serve HTTP on port 80. The controller and Agent Substrate both probe the container here. /healthz returns 200 as soon as the runner is alive; /readyz returns 503 until the workspace is prepared and 200 after; /metadata/v1alpha1/ax/task returns the Task spec as application/yaml; /metadata/v1alpha1/ax/workspaces returns every bound Workspace as a multi-document YAML stream.
  • Prepare each workspace once. For each binding, at its path: clone the Git repos from spec.git, create the skills path, write MCP config, run any goal-based bootstrap. A binding without a path lands at /workspace/<name>. Record that setup happened somewhere on the durable volume so later boots skip the work — “Resume restarts the container, and re-cloning into a restored workspace would destroy the agent’s state.” The default runner writes a marker file under /ax per workspace.
  • Run the command and supervise it. Start spec.command as a child with the first workspace as its working directory, give it AX_METADATA_URL plus every spec.env entry, put it in its own process group so you can signal everything it spawns.
  • Stay up after the command exits. The runner is PID 1. If the runner exits when the command does, the metadata server goes with it and ax ssh stops working. Log the exit status and keep serving until told to stop. The control plane does not currently read the command’s exit status back from the container — worth knowing if you build tooling around task completion.
  • Shut down cleanly on SIGTERM. Forward to the command’s process group, wait a bounded grace period, then SIGKILL whatever is left. Flush anything the agent needs to survive a resume before you exit.
  • Serve guest services only when asked. When spec.debug: true, the runner serves the Agent Substrate guest services over gRPC on port 80, multiplexed with HTTP via h2c. They are off by default because they allow arbitrary process execution and file access inside the sandbox. ax ssh refuses to connect to a task that hasn’t opted in.

The default runner image is Python 3.12 with git, curl, openssh-client, and the Antigravity agent installed, because the goal-based workspace bootstrap hands the goal to Antigravity. make build-task-runner cross-compiles for linux/amd64 and builds the image; make push-task-runner pushes it. The example Dockerfile in the runner doc pins to a specific image digest — gcr.io/ax-substrate/ate-images/ax-task-runner@sha256:69b764607ec7f1e433d83d2eca17dccfa04f663b43f071fd376e2dd716a57f8c — because the runner behavior has to be reproducible across builds.

There are three levels of customization the runner doc enumerates, in increasing order of commitment: extend the default image with RUN apt-get install and keep the entrypoint, embed the runner package in your own Go binary and call runner.Run directly, or write a runner from scratch that honors the contract. The first option is the cheap one; the second is the one a team running a specialized agent runtime will reach for. The example in runner.md shows the embedding path:

err := runner.Run(ctx, runner.Config{
    Task:       &task,
    Workspaces: workspaces,
    OnCommandExit: func(exit runner.CommandExit) {
        slog.Info("agent finished", "exitCode", exit.ExitCode)
        // Upload artifacts, notify a webhook, and so on.
    },
})

That OnCommandExit hook is the thing that lets you ship results out of a sandboxed task without the runner knowing about S3, Slack, or your webhook URLs. The runner stays domain-agnostic; the embedder plugs in the side effects. It’s a clean separation that a lot of agent runtimes skip — they tend to assume every team wants the same exit handling.

The goal-based workspace bootstrap is the closest thing to a “what is this framework actually for” tell

When a Workspace binding has a goal, the runner hands that goal to a bootstrap agent on first boot. The default image wires this to Antigravity, and the field the runner reads from the container is GEMINI_API_KEY — set automatically when the atespace has a Gemini credential configured through a Model resource. The task reports not-ready until every workspace, including any agent run, has finished. The default timeout is 10 minutes; set AX_BOOTSTRAP_TIMEOUT to a Go duration to change that.

This is the feature that, more than any other, tells you what Google thinks agents are going to spend their time doing. The goal is plain-language — “Install dependencies and run the test suite” is the example in examples/task.yaml. The runner hands it to an agent that finishes the setup. The implication: the framework assumes that getting an agent to a point where it can do useful work is the expensive part, and that “useful work” often starts with environment preparation. The agent doesn’t ship a fixed environment; it ships the description of one and lets another agent negotiate the steps.

The trade-off is the obvious one: you now need an agent capable of finishing the setup, you need to trust the agent with the right credentials, and you have given up a chunk of reproducibility in exchange for flexibility. A team that wants deterministic workspace setup can leave goal empty and rely on the Git clone + MCP + skills path skeleton. A team that wants “fresh container, fully-prepared workspace every time, paid for in agent runtime” sets a goal. AX doesn’t enforce one or the other.

Networking: the atenet router and the ate-target-actor header

The networking doc is short, which is a relief. Tasks do not get a Kubernetes Service or Ingress of their own. Every request goes through Agent Substrate’s atenet router, the atenet-router Service in the ate-system namespace. The router reads a single header, ate-target-actor, resolves the actor to the worker it is running on, resumes it first if it was suspended, and proxies the request there. The header value is <atespace>/<task>.

From inside the cluster, the canonical reach is:

curl -H "ate-target-actor: default/task123" \
  http://atenet-router.ate-system.svc.cluster.local/metadata/v1alpha1/ax/task

That is exactly how the controller polls a task’s readiness. From a laptop, you kubectl -n ate-system port-forward svc/atenet-router 8001:80 and add the same header. The Host and :authority headers are left alone for the application; only ate-target-actor selects the target. For gRPC, the header is outgoing metadata under the lowercase key.

The router-then-header pattern is what makes suspend/resume feel normal instead of magic. When a task is suspended, the actor is checkpointed; the router still knows the mapping default/task123 → worker. When the request comes in, the router resumes the actor first and then proxies. The header stays the same; the underlying worker may be a fresh container. ax ssh uses the same path.

Lifecycle in practice: what ax actually shows you

The CLI surface is wider than the four primitives suggest because most operational questions need more than the resource’s metadata:

ax apply -f examples/task.yaml       # Task + Workspace + Gateway + Model in one file
ax get tasks
# NAME      ATESPACE   PHASE     ACTOR           WORKER-IP    AGE
# task123   default    Running   task123         10.20.3.67   1m

ax watch task task123                # stream phase and condition changes live
ax ssh task123 -- ls -la /workspace  # poke around inside the sandbox
ax suspend task task123              # checkpoint actor state and pause
ax resume task task123               # pick up where it left off

The ax get tasks output columns are NAME, ATESPACE, PHASE, ACTOR, WORKER-IP, AGE — and WORKER-IP is the giveaway that the AX control plane is observing Agent Substrate workers directly. It’s not abstract; it’s a real IP on a node, and ax ssh will reach the sandbox through the atenet router using the ate-target-actor header that combines the atespace and task name.

The watch path is the operator’s primary tool during long agent runs. ax watch streams status.phase and condition transitions: WorkspaceReady flips to True once the bootstrap agent has finished; GatewayReady flips True once the egress policy has been applied; Ready is the one to wait on. Suspending a task sets Ready to False with reason TaskSuspended; resuming sets it back. Deleting a task moves it to Terminating while the controller tears down the sandbox, then removes the record entirely. ax delete blocks until that has happened.

The exit-code detail is worth restating because it surprised me: the runner is PID 1, the runner stays up after spec.command exits, and the control plane does not currently read the command’s exit status back from the container. If you build anything around “did this task succeed,” you have to instrument the runner yourself — typically through the OnCommandExit hook in the embedded-runner path. The control plane knows about lifecycle, not about success.

Trade-offs and what it doesn’t fix

Three honest limits are worth flagging for anyone evaluating AX against other agent infrastructure in 2026.

First, the framework is still moving fast under a “we may break things” warning. The repo is at v0.3.0, the most recent commit on main is a restructure that renamed the project’s self-description to “a general-purpose orchestration layer for agentic tasks” — a deliberate narrowing away from whatever the project was first scoped as. The release cadence (six releases since 2026-05-20: v0.1.0, v0.2.0, v0.2.1, v0.2.2, v0.2.3, v0.3.0) is steady, and the dependency bumps in mid-August suggest the maintainers are tracking upstream Go and Kubernetes API breakage, but there is no v1.0 commitment and the README’s WARNING block is the polite way of saying “the schema can move under you.” A team wiring AX into a production control plane should pin to a tag, not main, and budget for a migration cost.

Second, the runner is deliberately not the orchestrator. AX does not try to make decisions about which LLM an agent should call, how to budget tokens across a fan-out, or when to abandon a sub-task that is going sideways. The framework assumes those decisions live in your agent code, in the Workspace goal, or in the Model resource configuration. If you want an opinionated “use cheap model for planning, frontier model for code generation” policy, AX does not ship one. You write it yourself or wrap it in a separate layer.

Third, the egress policy is host-and-port, not identity-based. A Gateway allowlist says “this sandbox may reach api.anthropic.com on 443.” It does not say “this sandbox may reach api.anthropic.com as user X” or “this sandbox may call the messages endpoint but not the files endpoint.” The agent’s runtime has to enforce that. For teams that need per-endpoint, per-method authorization, AX is the orchestration layer that hands you a fenced network; the per-call guard is your responsibility.

Where it lands in the 2026 agent infrastructure landscape

The reason I sat with the repo for an evening rather than skimming the README is that the design choices line up with what Google has been building toward for agentic infrastructure in 2026. AX is not the runtime — that role is played by Agent Substrate, which AX sits on top of. AX is the control plane for the runtime: it decides what gets scheduled, what network the schedule runs inside, what model configuration the agent should pick, and what state persists across suspend/resume. The Redis-Streams-not-etcd choice is the most consequential design call, because it makes AX scale with task throughput rather than with object count.

The other 2026 agent orchestrators I’ve covered recently are mostly single-machine (Orca, the worktree-native ADE) or cluster-shaped but opinionated (Agent Substrate itself). AX occupies a narrower band: “Kubernetes-shaped, but for tasks that don’t want to be CRDs, on a runtime that is itself an open-source project.” The kubectl-shaped CLI and the four-primitive surface are the parts that make it approachable; the runner contract and the Redis Streams design are the parts that make it interesting.

Whether v0.3.0 holds long enough for production use is the open question. The honest version of the README’s WARNING block is: the maintainers are committed to the shape (Task, Workspace, Gateway, Model, atespace, runner) and uncommitted to the schema. The shape is what you write integration code against; the schema is what your YAML references. Pin to a tag.

References and where to dig further

  • Repo: github.com/google/ax
  • Design rationale: DESIGN.md — Redis Streams vs etcd, control-plane topology, four-binary architecture, full gRPC API reference
  • Core concepts: docs/concepts.md — Task lifecycle conditions, Workspace goal semantics, Gateway egress model, Model resource behavior
  • Manifest reference: docs/manifests.md — annotated Task, Workspace, Gateway, Model YAML with both Google and Anthropic provider examples
  • Runner contract: docs/runner.md — what a runner must do, three customization levels, Go embed example with OnCommandExit hook
  • Sandbox contract: docs/sandbox.md — PID-1 runner responsibilities, metadata server endpoints, guest services gating, environment variables
  • Networking: docs/networking.md — atenet router, ate-target-actor header, in-cluster and laptop access patterns, gRPC metadata key
  • Example manifests: examples/task.yaml, examples/multi-workspace.yaml — multi-document YAML with one of each primitive
  • Underlying runtime: github.com/agent-substrate/substrate — the sandboxed execution layer AX runs on
  • Guest services spec: github.com/agent-substrate/env — what ax ssh and spec.debug: true expose
  • CLI tooling dependencies: ko for control plane image builds, kubectx for cluster switching that ax --context follows
Aniket Karne
DevOps & AI Engineer · Amsterdam
Back to all posts
Reader correspondence

Comments

Powered by GitHub Discussions via Giscus. Sign in with GitHub to leave a comment.