TraceCrate: The Privacy-First Agent Trace Workbench That Refuses to Be a Platform — aniketkarneai.com | aniketkarneai.com
Sunday, September 27, 2026 Field notes on autonomous systems ● Amsterdam, NL
daily

TraceCrate: The Privacy-First Agent Trace Workbench That Refuses to Be a Platform

TraceCrate v0.2.0 reads Claude Code, Codex, and OTLP traces entirely in the browser — no backend, no telemetry, hard resource caps, structure-only export with optional metadata minimization, and jsdiff comparison that returns an explicit unavailable state when a budget is exceeded instead of a misleading partial result.

Most agent observability tools look the same way: a hosted dashboard, an SDK that ships your traces to a SaaS endpoint, a billing tier. TraceCrate, which hit v0.2.0 today at dc5d92daabe84378d994f09637db317f36b21024, is the opposite shape. It is a static Vite-built SPA that runs in the browser, parses JSONL transcripts from Claude Code and Codex and OTLP resourceSpans payloads, and never opens a network socket after the page loads. There is no backend. There is no collector. The release notes are unusually explicit: “the app has no trace-upload or telemetry path. That does not make every environment or export safe.”

That second sentence is the design philosophy. TraceCrate treats privacy as a boundary it can document but not guarantee, and treats every other constraint in the same way. When the parser hits a 30-second deadline on a 20 MiB file, it cancels the worker and discards the pending batch. When the comparator runs over its 2,000-events-per-side or 200-millisecond jsdiff budget, it returns an explicit “unavailable” state instead of a misleading partial diff. When you ask for a structure-only export of a session, the redaction layer enforces an allowlist — original text, names, model IDs, and unknown properties are removed by default; timing and token usage stay in by default but can be omitted by toggling omitTiming and omitUsage. The default-and-override shape is documented in a table you can read end to end.

What makes TraceCrate worth a post is not that it exists — there are a half-dozen agent-trace viewers in various states of abandonment on GitHub. What makes it worth a post is the granularity of the design choices and the discipline of the v0.2 release. Let me walk through the parts that stuck out.

The pipeline, in stages

The architecture is documented as a six-stage table in docs/architecture.md. File selection and session state live in App.tsx and useTraceImport.ts. Background import runs in an inline module worker spawned per file — fresh worker per import, terminated on completion, error, cancellation, clear, unmount, or the 30-second deadline. The worker does UTF-8 size/depth/record checks before any parsing happens. If the file is over 20 MiB or 20,000 records, the worker rejects it before reading further. Decode/dispatch tries JSON first, then JSONL, runs an auto-discovery pass over registered adapters, and picks the first match. Adapters normalize the source-specific format into a common Trace / TraceEvent / Usage schema, which gets validated by Zod. Then analysis, query, comparison, and rendering run on the UI thread. Export runs on the UI thread too.

Two things in that pipeline matter more than they sound.

First, the worker model. A worker per file means a slow import on one trace cannot stall imports on the rest of a batch, and a malformed file cannot corrupt a worker’s internal state across imports. The worker removes its message/error/abort handlers and its timer exactly once. A generation token prevents late results from repopulating a cleared or cancelled workspace. This is the kind of hygiene that prevents a class of bugs you only find when somebody clears the workspace while an import is mid-flight — bugs that look like ghosts and trace to a stale worker reply.

Second, the analysis pass on the UI thread. The architecture document says this plainly: “Web Workers are required for UI imports; no synchronous fallback is provided. A worker protects responsiveness during parsing, not during all analysis/search/export work.” TraceCrate is honest that web workers do not save you from a 200-millisecond jsdiff alignment on the main thread. The budget is a guardrail, not a guarantee. The same paragraph says there is no guaranteed memory ceiling; input copies, normalization, main-thread analysis, comparison, and exports all consume memory, and the browser is still trusted to run scheduled callbacks.

The heuristics, by name

The findings module is the closest thing TraceCrate has to opinionated analysis, and the heuristics are listed verbatim in architecture.md: explicit tool errors; tool duration ≥10,000 ms; recorded output ≥12,000 JavaScript string units; at least 3 calls with exactly equal event names and input text. Each is a literal pattern match against the normalized trace, not an inference. Missing input is not evidence of identical input. Equivalent JSON with different whitespace is not equal text. Findings link to observed events, not explanations of intent.

getStats uses reported usage and timestamp endpoints. getToolBreakdown sums observed durations, which can overlap. Comparison is B − A for recorded values, not experimental control, not statistical inference, not scoring, not cost estimation. The built-in demo generator manufactures sequences, token counts, and timings; the “optimized” label on the demo comparison is illustrative only.

I appreciate that TraceCrate ships with a manufactured demo that says, out loud, “this is not a benchmark.” A surprising number of trace viewers ship with a demo that quietly implies its numbers mean something. TraceCrate treats that implication as a bug to design away.

The comparator and its budget

The v0.2 change that landed in compareEvents is the most interesting one. TraceCrate now aligns two recorded sessions event by event using jsdiff.diffArrays with exact JSON-encoded (kind, name) keys. Aligned events compare content, input, output, status, model, duration, and usage. IDs, absolute timestamps, and parent links are not cross-run identities — TraceCrate says so explicitly. Repeated names can align ambiguously, and the comparator says so explicitly. Insertions or removals mean present on that side of this alignment, not success or failure.

The budget is concrete: at most 2,000 selected events per side, 400 edits, 200 milliseconds alignment time. If any of those bounds is exceeded, the comparator returns an explicit unavailable state — not a partial diff, not a truncation warning, not a “best effort” message. The UI shows “unavailable” and stops.

This is a design choice you do not see in most trace viewers, which prefer to render something rather than admit they cannot finish. TraceCrate’s reasoning is in the design doc: “Exceeded size/edit/time bounds return an explicit unavailable state, never a misleading partial comparison. Work still occurs on the UI thread; the budget is a guardrail, not a strict latency guarantee.”

The export pipeline, with minimization

The structure-only export mode is the privacy story. By default, the allowlist removes original text, names, model names, inputs, outputs, and unknown properties. IDs and valid parent links are remapped to local random IDs. Token usage and timing stay in by default; toggling omitTiming or omitUsage removes them. Both flags are valid only in structure mode — pattern mode plus minimization throws, because pattern redaction preserves text and can miss secrets. The transformed JSON and HTML downloads are sharing transforms, not raw backups; the native schema stays at version 1.

The HTML export has no scripts and no external assets. The JSON export can be reimported. Omitted values are absent, never zero — which matters, because a zero in a downloaded report reads like evidence, and TraceCrate refuses to manufacture evidence.

The privacy document draws a clean line: “Structure-only removes arbitrary free text, original identifiers/names, inputs, outputs, and model names using a strict allowlist. Token usage and timing remain by default and can now be omitted; event order, counts, remapped relationships and statuses still remain and can be sensitive. Pattern redaction preserves text and can miss secrets. Neither mode guarantees anonymization; inspect every field of the full download. Clearing memory is not secure erasure.”

That is the second sentence I quoted at the top: privacy is a boundary, not a guarantee. The export dialog includes a “Show redacted preview” button that renders the exact bytes of the download before you save them, which is the operational counterpart of “inspect every field.”

Format support, by honest limits

TraceCrate reads three formats plus its own native schema. Claude Code: normal message JSONL — text, tool calls/results, selected system/result records. Partial streaming deltas are ignored. Private transcript variants can differ. Codex: rollout JSONL with session_meta, turn_context, response_item, and selected event_msg. Not arbitrary codex exec --json events; reasoning and deltas are not imported. OTLP JSON: resourceSpans → scopeSpans → spans with nested structured attributes, current and legacy cache-write keys. No protobuf, no collector endpoint, no full OTLP, no MCP transcript support. TraceCrate native: one validated schemaVersion: 1 JSON report, with unknown fields stripped.

The format doc is unusually precise about what is not supported. That precision matters because a viewer that silently ignores streaming deltas or strips reasoning traces will produce a comparison report that misses exactly the differences a careful engineer is looking for. The doc spells out the boundary so the user does not have to discover it by running a real session through.

The v0.2 release added OTLP kvlistValue parsing and the new gen_ai.usage.cache_write.input_tokens name from the OpenTelemetry GenAI semantic conventions, with a legacy fallback for the older cache-creation name. Both names mean the same thing; the rename is the kind of break that bit me on an OpenLLMetry export a few months ago, and TraceCrate handles it without ceremony.

The release process and the HTTP 500 recovery

The v0.2.0 release notes describe a release validation job that passed 225 unit tests and all 64 browser tests, plus lint, application/browser TypeScript, build, coverage, and dependency audit. Core/adapters line coverage is 97.51% — explicitly not UI coverage. The publication was recovered from an HTTP 500 using the unchanged validated artifact, and the overall workflow remains marked failed, not passed. Both public downloads were independently verified against SHA-256 and the original artifact.

That is a remarkably honest release note. The CI workflow itself failed; the artifact was good; the release was published anyway from the artifact, and the failure is documented in the run notes rather than hidden behind a green badge. I have shipped releases that look exactly like this — green tests, red workflow, “is the artifact safe to publish” — and the right answer is “yes, publish from the artifact, document the workflow failure, fix the workflow in the next commit.” TraceCrate did that.

Trade-offs and what it doesn’t fix

Three honest limits.

First, the in-memory workspace caps at 10 sessions including the two demos. That is enough for one or two active debugging sessions, not enough for a team’s worth of historical traces. If your debugging flow involves scrolling through fifty traces from the past week, TraceCrate is not the right tool. The release does not claim otherwise.

Second, the analysis pass runs on the UI thread. The architecture doc says so plainly. A 2,000-events-per-side comparison at the 200-millisecond budget will block input on a slow laptop. TraceCrate’s answer is “the budget is a guardrail” — which is a true statement, but it does mean the tool can stall on a comparison that would be smooth on Langfuse. The trade-off is the privacy story: no backend means no server-side alignment.

Third, the structure-only export removes arbitrary free text using an allowlist. That removes secrets that look like arbitrary free text. It does not remove secrets that are structured — for example, an API key embedded in a tool argument that gets serialized as a JSON string is exactly the kind of field the allowlist is designed to remove, but a secret encoded as a tool name or a model parameter is not. TraceCrate says so: “Pattern redaction preserves text and can miss secrets.” Use structure-only and inspect every field.

The thing I am still sitting with is the comparator’s “unavailable” state. It is the right answer for an honest tool — admit when you cannot finish rather than render a misleading partial result. But it also means TraceCrate punts on the very comparison a serious postmortem often needs: what changed between the session that worked and the session that didn’t, when the change was a long agentic loop with hundreds of tool calls. The 2,000-events-per-side cap will hit that case. I do not know what the right answer is, and TraceCrate does not pretend to know either.

If you want to try it without installing anything, the live demo at fankchen.github.io/tracecrate loads two synthetic runs and walks you through Timeline, Insights, Compare, and Export. The v0.2.0 release and the downloads are at github.com/FankChen/tracecrate/releases/tag/v0.2.0. The native schema is version 1, the source is MIT-licensed, and the project is not affiliated with Anthropic, OpenAI, or OpenTelemetry.

Aniket Karne
DevOps & AI Engineer · Amsterdam
Back to all posts
Reader correspondence

Comments

Powered by GitHub Discussions via Giscus. Sign in with GitHub to leave a comment.