Paperclip: The Company Around Your AI Agents — A Control Plane That Treats Heartbeats as Database Rows — aniketkarneai.com | aniketkarneai.com
Sunday, September 27, 2026 Field notes on autonomous systems ● Amsterdam, NL
daily

Paperclip: The Company Around Your AI Agents — A Control Plane That Treats Heartbeats as Database Rows

paperclipai/paperclip (87k★, MIT, v2026.916.1) is building the HR system for fleets of AI agents — Companies, Boards, cascading budgets, durable heartbeats as DB rows, and a separate Rust Paperclip Runner Protocol for executing Codex/OpenCode/Claude Managed/AWS AgentCore behind a default-off flag.

If you’ve ever had 20 Claude Code tabs open at once and lost track of which one was doing what, the README of paperclipai/paperclip starts with a sentence that hits: “If OpenClaw is an employee, Paperclip is the company.” It then spends the next 29 kilobytes describing what that means in practice. The repo crossed 87k stars this week, ships a new version roughly every two weeks (current: v2026.916.1, Sep 21 — six days old at this writing), and quietly became the most ambitious multi-agent orchestration project I’ve read this month. It is not an inference engine, not a memory layer, not a prompt registry. It is a control plane for running AI agents as if they were employees inside an actual company — with org charts, budgets, board oversight, approval gates, and a heartbeat system that survives process restarts because the heartbeat is not in memory, it is a database row.

This post is what I learned reading the doc/SPEC.md (35kB), the four documents under doc/architecture/, and the v2026.916.x release notes for MCP Access Governance. There is a lot of surface here. I’ll focus on the three pieces that surprised me most: the Paperclip Runner Protocol as a separate Rust process behind a default-off flag, the native status arbitration pipeline that makes “the model said it was done” not authoritative, and the durable continuation scheduler that keeps heartbeats going without keeping an agent process alive.

What “the company around the agents” actually means

Paperclip’s first-order object is a Company, not an agent. One Paperclip server runs multiple Companies. A Company does not have a single goal field — direction is encoded as the set of Initiatives the company is running. Every Company has a Board: the human oversight layer. In v1 the Board is a single human operator, and the Board has unrestricted access — set budgets, pause any agent, pause any work item, override any agent decision, manually change any budget at any level. The Board is not just an approval gate; it is a live control surface the human can intervene on at any level at any time.

Below the Board sit Agents, each with an adapter type and an adapter-specific config blob. The adapter defines what an agent is — OpenClaw adapter uses SOUL.md and HEARTBEAT.md, Claude Code adapter uses CLAUDE.md, HTTP adapter takes an endpoint URL + auth headers + a payload template. Paperclip does not prescribe how an agent defines its identity or behavior; it provides the control plane and lets the adapter define the agent’s inner workings. The org chart defines reporting and delegation lines; visibility is full-everywhere by default in v1. Work-object privacy is explicitly not a v1 feature until centralized scoped authorization lands.

This is the unusual move. Most multi-agent frameworks I read in September are bottom-up: they expose primitives (tools, hooks, memory, retries) and let you assemble a team. Paperclip is top-down: it ships a Company abstraction with a Board and assumes you are operating a company, with all the governance baggage that implies. The README’s “Problems Paperclip solves” table makes the bet explicit: 20 Claude Code tabs, manual context-gathering, folder-of-configs disorganization, runaway loops, recurring jobs you have to remember to kick off, finding-the-repo-and-babysitting-the-tab. The thesis is that if those problems are what you actually have, then the abstraction you need is not “agent with memory” or “agent with retries” — it’s “company with HR.”

Two interesting design choices in how that plays out:

  1. Three integration levels. Paperclip requires only that an agent be callable (Paperclip can start it via command or webhook). Status reporting is optional. Fully instrumented — reporting status, cost/token usage, task updates, logs — is the third tier, and the default agents shipped in the repo demonstrate the full integration as reference implementations for adapter authors. This is a deliberately low floor. The “Bring Your Own Agent” table shows OpenClaw, Claude Code, Codex, Cursor, Bash, HTTP — anything that can receive a heartbeat.
  2. Two context-delivery modes. A fat payload bundles relevant context (current tasks, messages, company state, metrics) into the heartbeat invocation — suited for stateless agents that cannot call back. A thin ping is just a wake-up signal; the agent calls Paperclip’s API to fetch what it needs — suited for sophisticated agents that manage their own state. The choice is per agent. It means a heterogeneous fleet can use both shapes without forcing one adapter model on everyone.

There is also a portability story I want to flag: the entire org’s agent configurations are exportable. Template export ships structure (agent definitions, org chart, adapter configs, role descriptions) for spinning up a new company. Snapshot export includes current tasks, progress, and agent status. Secret scrubbing and collision handling are first-class. You can fork a marketing-agency company template into a different Company in the same Paperclip instance.

Paperclip Runner (PRP) — a separate Rust process behind a flag

The single most interesting architectural decision in Paperclip right now is the experimental Paperclip Runner (paperclip-runnerd), a standalone Rust process that the server opens a native run against. The ADR (doc/architecture/paperclip-runner.md, dated 2026-08-24) is explicit about what the runner is not: it is not a second control plane, it does not own identity, it does not own policy, it does not own workflow state. The server is the authority for all of those. The runner owns durable delivery, restart recovery, and governed access to Paperclip actions for one provider session. Its qualified provider catalog: Codex, OpenCode, Claude Managed, AWS AgentCore, plus pinned Claude/Codex ACPX profiles.

The topology is:

Paperclip server
  |  authenticated PRP v1 WebSocket
  v
paperclip-runnerd
  |  qualified native provider protocol
  v
Codex / OpenCode / ACPX / Claude Managed / AWS AgentCore

The WebSocket is at /api/runner/v1/connect/:runId. Runnerd opens the outbound connection; the server launches a verified runnerd artifact in the realized execution environment. The browser never connects to runnerd — it reads projections from the existing Paperclip APIs and task-thread models. This split lets the server keep the policies and identity while pushing the volatile provider process out into its own lifecycle.

Two things from the ADR are worth highlighting:

  • Deterministic parity TS↔Rust. The dependency direction is implementation → contract, not the reverse. The runner package must build and test without importing Paperclip server, UI, CLI, database, or other private workspace implementation modules. There is a JSON Schema and fixtures package at the bottom of the dependency tree that both TypeScript contracts and the Rust runner core consume; a deterministic parity package verifies both sides agree. The stated goal — “make protocol behavior deterministic across TypeScript and Rust” — is the kind of thing that gets expensive if you don’t enforce the dependency direction from day one, and the ADR is explicit about it.
  • Default-off rollout flag. The paperclip_runner adapter is only available when an instance-level, default-off rollout flag is enabled. Existing direct adapters keep their current invocation, transcript, interaction, cancellation, and finalization paths. The compatibility invariants (doc/architecture/paperclip-runner-compatibility.md) are ten numbered rules, including “a non-runner run must not start runnerd or open PRP,” “a non-runner run must not create native result records,” and “recovery may finish an already persisted native run while fresh native starts remain blocked.” The persisted-runtime table is explicit: the server resolves and persists the runtime once, before provider launch, and never silently falls back from paperclip_runner to codex_local. A configuration error must be visible, not hidden.

The Codex compatibility window is also pinned to a real range: remote native Codex runs accept stable CLI versions >=0.149.0 <0.157.0, with install pin 0.156.0 (recorded on 2026-09-22). The minimum is fixed at 0.149.0 until maintainers deliberately change it — not a rolling one-month support window.

Native status arbitration — model prose is not a privileged status command

This is the part of the architecture that I think is genuinely underrated. The runner returns a paperclip.run_result.v1 result with fields like reportedWorkDisposition (done / blocked / needs_review / yielded), completionClaim (with contract revision, criterion claims, and remaining work), verification claims, evidence references, optional blocker, optional continuation. The pipeline is then:

structured runner result
        |
        v
schema and terminal validation
        |
        v
evidence classification
        |
        v
pure status arbitration
        |
        v
transactional decision commit
        +--> issue status/version
        +--> durable side effects
        +--> audit and recovery records

The reason this exists is exactly what the doc says: “This separation prevents model prose from acting as a privileged status command, protects newer issue state from stale runs, and makes every decision replayable and auditable.” In other words, the model can report that work is done, but the server is the authority that decides and commits the resulting workflow state. classifyNativeEvidence() compares model claims against durable Paperclip records and recognizes specific evidence families — the server owns facts the runner cannot choose (the persisted completion contract, the run’s actual terminal state, whether workspace finalization succeeded, the issue’s current status and status version, pending approvals, the completion-authority policy recorded on the contract).

If you have ever watched an agent say “I’m done” and then discovered the diff is empty, this is the abstraction that fixes it. The agent reports claims; the server arbitrates against the actual record; the decision is a transaction, not a model output. It also means an issue’s status is owned by the same control plane that owns the rest of the workflow — there is no separate “agent-state vs server-state” drift.

Durable continuation — heartbeats as database rows

The third piece of the architecture is doc/architecture/durable-continuation-scheduler.md. The core claim: Paperclip does not keep an agent process alive between turns. A heartbeat run is finite — it starts, performs work, records a terminal result, and exits. If work must continue later, Paperclip represents that intent in database state and creates another heartbeat run when the continuation becomes eligible.

Two mechanisms provide this:

  1. Explicit continuation effects — a native result can report yielded with a continuation (kind: same_agent, retry, delegated_issue, response_wake, or monitor) plus a summary and an idempotency key. The native status arbiter keeps the issue in_progress and emits an enqueue_continuation effect; the status-decision committer materializes it as an idempotent agent_wakeup_requests row.
  2. Stranded-issue reconciliation — reconcileStrandedAssignedIssues() is a startup and periodic safety net that scans agent-owned issues in todo, in_progress, and relevant in_review states. The scheduler runs unless HEARTBEAT_SCHEDULER_ENABLED=false; the interval is configured via HEARTBEAT_SCHEDULER_INTERVAL_MS (default 30,000 ms, clamped to a minimum of 10,000 ms).

The system reconstructs intent from persisted control-plane records rather than an in-memory timer owned by an agent. The records that survive a process restart are: the issue status and assignee, the issue execution lock/run identity, heartbeat run status/context/retry ancestry/terminal timestamps, agent_wakeup_requests rows, native status decisions and their materialized effects, pending interactions/approvals/monitors/blockers/execution stages, scheduled retry timestamps, and recovery actions.

The contrast to “agent with a long-lived process” is the point. A long-lived process gives you statefulness for free but loses you restart recovery, audit, and the ability to do native status arbitration. By making the heartbeat finite and persisting its continuation as a row, Paperclip gets restart recovery, audit, and the ability to gate every state transition through the server. The trade-off is operational: every continuation has to wait at most one scheduler interval before it shows up as a new run, and the reconciliation is polling rather than event-driven at the seam.

MCP Access Governance — the v1 release that’s actually about trust

The most recent substantive release is v2026.916.1 (Sep 21) with the MCP Access Governance feature (doc/RELEASE-NOTES-mcp-access-governance.md, source issue PAP-10397). The TL;DR: Paperclip now governs every MCP tool call an agent makes. Operators install managed connections, define profiles and policies, approve high-risk actions, and read an append-only audit log. Default posture: deny on unknown tools; quarantine on schema drift; approval required by default for writes; trust rules to lift human-in-the-loop on safe repeats.

The conceptual model is a layered one:

  • Connections — remote_http (preferred, hosted SaaS MCP) or local_stdio (approved-template-only, gated by deployment trust). Operators never paste raw stdio commands; templates ship in the build.
  • Catalog with risk classification — every discovered tool gets read / write / destructive inferred from MCP annotations. Destructive tools and unexpected new write tools are auto-quarantined.
  • Profiles + bindings — named bundles of include/exclude entries over the catalog, bound to a company, agent, project, routine, or issue. Narrowest binding wins.
  • Policies — allow, block, require_approval, rate_limit, trust_rule. Deny beats allow. Policies stack with profiles and run in priority order.
  • Trust rules — promote an approval into a scoped allow rule tied to the canonical argument hash and the catalog schema hash. When schemas drift, trust rules stop applying and the gateway falls back to approval. Revocations are first-class and audited.
  • Runtime supervisor — stdio runtime slots have a real lifecycle (starting, running, idle, failed), idle eviction, restart suppression on storms, and a board health endpoint with alert recommendations.

The default posture is deliberately conservative. An unknown tool is denied. Catalog drift (a new write or destructive tool seen on a refresh) is quarantined. A write tool with no policy match and no trust rule requires approval if the profile’s default-action allows writes; otherwise it is denied. A destructive tool is denied until an operator explicitly un-quarantines it. Local stdio in authenticated/public fails closed unless PAPERCLIP_TRUSTED_MCP_RUNTIME_HOST is set on a designated trusted worker. Agent-supplied stdio commands are rejected, always.

For agents the contract is simpler: agents speak MCP only to the Paperclip gateway. The gateway returns the agent’s effective tool list, validates each call against profile + policies, and records the result. When a call resolves to require_approval, the agent’s call blocks until a human decides. The agent does not see why a call was denied — that detail is in the audit log for the operator, on purpose, because “agents must not learn to route around denials.” Tool call arguments and results are subject to a redaction plan recorded on each call event.

The audit ledger is append-only: decision, matched policy IDs, reason code, redaction plan, latency, outcome. Per-run timelines are exposed at …/runs/:runId/decisions. Audit-write failures use a durable runtime metric counter — the release notes flag a firing mcp_runtime_audit_write_failures alert as a control-plane incident until audit durability is restored. That’s a small detail but it tells you the audit log is treated as a control-plane resource, not a debugging convenience.

A safe-read-only-todo-kv example bundle installs an application, a connection against a synthetic local fixture, and a read-only profile in one call. The bundled smoke check exercises allow_read_tool / deny_write_tool / audit_written and returns ok: true across all three — a self-contained way to confirm the gateway is healthy without an upstream MCP dependency. For an operator who has never used MCP Access Governance before, this is the kind of thing that earns trust fast.

Trade-offs and what it doesn’t fix

Three honest limits:

  1. Visibility is full-everywhere in v1. Work-object privacy is explicitly not a v1 feature until centralized scoped authorization is in place. For an early-stage team with one human Board this is fine. For a company that needs to run multiple Boards with strict data isolation between them, the only isolation primitive in v1 is Company-level (one deployment runs many Companies; complete data isolation is enforced there). Within a Company, every agent sees everything.
  2. The runner is experimental and qualified. Paperclip Runner is behind a default-off flag, the provider catalog is qualified (Codex / OpenCode / Claude Managed / AWS AgentCore plus pinned ACPX profiles), and the Codex CLI window is >=0.149.0 <0.157.0. If your provider isn’t in that catalog, the runner adapter won’t accept it. The direct adapters (which pre-date the runner) remain authoritative and unchanged — Paperclip is explicit that you do not route existing adapters through the runner.
  3. MCP Access Governance has a fixed expiry for action requests and no bulk catalog review. Action request expiry is fixed by policy; approvers cannot extend from the UI. Each quarantined entry is reviewed one at a time (no bulk review). Trust rules match exact argument shapes only — pattern-based trust rules are post-v1. Rate limits are per-policy, not cross-policy aggregates. None of these are fatal, but they are real friction points if you are evaluating this for production.

There’s also a thing the docs flag that I want to be honest about: the endpoint mode — Paperclip’s own /mcp surface — is not subject to the profile/policy stack. If you expose a Paperclip endpoint to other agents, the tool access stack does not gate it. This is the documented exception; it’s not a bug, but it’s worth knowing before you wire Paperclip into a multi-tenant gateway.

The thing I’d want next — and it’s not a criticism so much as a “watch this” — is what happens when the runner and the governance stack meet. The MCP Access Governance is the v1 trust layer; the runner is the v1 execution layer; the arbitration layer is what sits between them. The release notes for v2026.916.x don’t yet describe how a runner session reports a tool call that needs approval — whether the runner blocks waiting for the action-request to resolve, whether the server schedules a continuation, what the audit row looks like. That’s a real seam, and I’d expect the answer to land in the next couple of releases.

The bet, in one sentence

Paperclip is betting that the bottleneck for production AI agent teams in 2026 is not better models or faster inference — it is the missing organization. The control plane that ships with Paperclip — Companies, Boards, cascading budgets, durable heartbeats as DB rows, native status arbitration that makes model prose non-privileged, and now MCP Access Governance — is the answer to that bet, and the answer is unusually opinionated for a project that ships as MIT. Whether the bet pays off depends on whether teams with five-to-fifty agents end up needing a Company abstraction more than they need a smarter primitive. The repo’s star count and release cadence suggest the answer is already yes for a non-trivial slice of the early adopter base.

If you want to dig further, start with doc/SPEC.md (35kB, the canonical spec), doc/architecture/paperclip-runner.md (the runner ADR, dated 2026-08-24), and doc/architecture/durable-continuation-scheduler.md. The release notes for v2026.916.x are at doc/RELEASE-NOTES-mcp-access-governance.md and the operator-facing concepts are in doc/MCP-ACCESS-GOVERNANCE.md. The MCP demo script in doc/MCP-DEMO-SCRIPT.md walks the bundled safe-read-only-todo-kv smoke check end-to-end. The announcements/current.json and announcements/examples/ directories carry the canonical release summary if you want the short form.

Aniket Karne
DevOps & AI Engineer · Amsterdam
Back to all posts
Reader correspondence

Comments

Powered by GitHub Discussions via Giscus. Sign in with GitHub to leave a comment.