If you shipped an AI agent pipeline in 2024, the observability story was mostly “log the prompt and the response, hope for the best.” In 2026, the tooling has caught up. There are now real tracing platforms, real evaluation platforms, and a clearer picture of what the security surface looks like. This post walks through the actual landscape — the tools, what they do, and where they fit in a working agent stack.
This isn’t a vendor comparison with a “winner.” It’s a map of the territory, organized by what problem each tool solves.
The two-axis tool landscape
The 2026 AI observability space has settled into a two-axis structure:
Axis 1: tracing vs evaluation. Tracing platforms capture every span of an agent run — the prompt, the LLM call, the tool call, the subagent dispatch, the final response — and let you replay it. Evaluation platforms attach quality signals to traces — was the answer correct, was it grounded in context, did it leak sensitive data, did the user accept it. The two overlap but the architectural center of gravity is different.
Axis 2: open-source vs hosted. Open-source tools (Langfuse, Arize Phoenix, MLflow, Opik) can be self-hosted with no data leaving your infrastructure. Hosted tools (LangSmith, Confident AI, LangWatch) offer less operational burden but require sending traces to a third party. For data-residency-bound workloads, the open-source tier is the only option.
Most serious pipelines end up running two tools: one tracing platform and one evaluation platform, often from different vendors. The evaluation platforms import traces from the tracing platforms via OpenTelemetry, so the integration cost is low.
The tracing platforms
Langfuse
License: MIT (open-source core). Langfuse is the most-deployed open-source tracing platform in 2026. The architecture is OpenTelemetry-native, which means traces from any OTel-instrumented framework (OpenAI Agents SDK, Anthropic SDK, LangChain, custom code) flow into Langfuse without per-vendor adapters.
What it gives you:
- Per-request trace capture: every LLM call with prompt, response, model, tokens, latency, cost
- Multi-turn conversation tracking: session-level rollups of token spend and latency
- Tool call visibility: which tools were invoked, with arguments and results, in what order
- Subagent dispatch tracking: parent-child span structure for agent orchestration
- Self-hosted: Docker Compose, Kubernetes, or the managed cloud offering
- Eval integration: Langfuse has its own evaluation surface, but also integrates with external eval platforms
For a team that needs tracing and is unwilling to send data to a hosted vendor, Langfuse is the default choice. The MIT license is permissive; the operational burden is moderate (one Docker container plus a Postgres plus a clickhouse).
Arize Phoenix
License: Elastic License 2.0. Phoenix is the open-source tracing and evaluation platform from Arize AI. The Elastic License is non-OSI but permits self-hosting and commercial use with restrictions on offering Phoenix-as-a-service. For most teams that’s not a constraint.
Phoenix differentiates on evaluation depth:
- LLM-as-judge evaluation with prompt-template management
- Retrieval-grounded QA evals (RAGAS-style)
- Span-level attribution: which span in the trace is responsible for which quality issue
- Drift detection: statistical monitoring of embedding distributions to detect when the model’s input distribution shifts
If your agent pipeline includes RAG or retrieval-heavy workflows, Phoenix’s retrieval evals are the strongest in the open-source tier. The drift-detection infrastructure is also distinctive — most tracing platforms don’t do distribution monitoring.
LangSmith
License: proprietary, hosted only. LangSmith is LangChain’s tracing platform. It’s the most polished tracing UX in the category and the deepest LangChain integration. If you’re building on LangChain or LangGraph, LangSmith is the path of least resistance.
What it gives you:
- Polished trace explorer with timeline visualization, span detail, and replay
- LangGraph-native orchestration: traces map cleanly to LangGraph nodes
- Dataset management for evaluation: production traces can be promoted to eval datasets
- Hub for prompt templates: versioned, tested, deployable prompts
The trade-off: hosted only. LangChain has not open-sourced the core tracing platform. If you’re building on LangChain and willing to send traces to LangChain’s infrastructure, LangSmith is the smoothest experience. If you need self-hosting, look at Langfuse or Phoenix.
MLflow
License: Apache 2.0. MLflow is the traditional ML platform that has expanded into LLM tracing. The LLM tracing features are recent (the mlflow.genai namespace), and the platform retains the traditional ML strengths — experiment tracking, model registry, deployment.
What it gives you:
- Unified tracking: classical ML experiments and LLM traces in the same UI
- Model registry: deploy LLM artifacts alongside traditional ML models
- Self-hosted with Apache 2.0: most permissive license in the category
- MLflow evaluation: traditional ML metrics extended to LLM scenarios
For a team that already runs MLflow for classical ML, adding LLM tracing is a one-package install. For a team starting fresh on AI observability, MLflow’s LLM-specific UX is less polished than Langfuse or Phoenix.
Opik
License: Apache 2.0 (from Comet). Opik is the open-source tracing and evaluation platform from Comet. It’s newer than Langfuse or Phoenix and has more polish on the evaluation side than the tracing side.
What it gives you:
- Tracing with span-level visibility
- Built-in evaluation suite: heuristics, LLM-as-judge, custom scorers
- Experiment comparison: side-by-side runs of the same agent with different prompts or models
- Apache 2.0 license: fully permissive, including the offering-as-service restriction that’s missing from Phoenix
For a team that wants evaluation depth and a permissive license, Opik is worth evaluating. The community adoption is smaller than Langfuse, which means fewer integrations but also less vendor lock-in.
Helicone
License: MIT. Helicone is the lightweight proxy-based tracing platform. You route your LLM API calls through a Helicone proxy and the proxy captures every request automatically. The integration cost is zero — no code changes, just a base URL swap.
What it gives you:
- Drop-in observability for OpenAI-compatible APIs (most LLM providers)
- Per-request cost tracking with model-aware pricing
- Caching: identical requests can be served from cache, reducing API spend
- Prompt versioning: track prompt variants through the proxy
For teams that want quick observability without instrumenting their application code, Helicone is the lowest-friction option. The proxy architecture means less application-level context (no subagent tracking, no tool-call visibility beyond what the LLM API returns), but the cost and latency tracking is solid.
The security layer
The observability story isn’t complete without the security story. Three threat surfaces matter for AI agents:
1. Prompt injection
Prompt injection is the structural vulnerability where untrusted input (web content, retrieved documents, tool outputs) gets concatenated into the model’s context and overrides the system prompt. The OWASP LLM Top 10 lists it as the #1 risk.
Mitigations:
- Prompt hierarchy with role separation: system prompt, developer prompt, user prompt as distinct message roles; instruct the model to never treat retrieved content as instructions
- Instruction-data separation: tag retrieved content as
<data>blocks with explicit non-instruction framing - Output validation: treat model output as untrusted; require structured outputs that can be validated
- Tool-call authorization: every tool call must be authorized against an explicit allowlist, not derived from the model’s reasoning
- Dual-LLM pattern: use a “privileged” LLM and an “untrusted” LLM in tandem; the untrusted LLM sees external content, the privileged LLM only sees structured outputs
None of these are perfect. Prompt injection is fundamentally a context-confusion attack and the mitigations are partial. The operational discipline is “treat every external content source as adversarial, validate everything.”
2. Secret leakage
Agents that have access to credentials (API keys, database URLs, OAuth tokens) can leak those credentials through output. The risk is highest when the model is asked to summarize, log, or display its context.
Mitigations:
- Secret redaction in observability: tracing platforms must redact secrets before logging. Langfuse, Phoenix, and LangSmith all support redaction patterns; verify your platform is configured to redact the specific secret formats you use.
- Credential scoping per tool: don’t give an agent a “do everything” credential. Each tool should have a credential scoped to its specific action.
- Masked credential request path: when an agent needs a credential mid-run, the request should go through a masked channel (like OpenClaw’s #129670 or equivalent). The credential value should never enter the model’s context.
- Egress allowlisting: outbound network calls from the agent runtime should be allowlisted to known-good destinations.
3. Tool-call abuse
Even with a well-instructed agent, the model can be tricked into making tool calls with attacker-chosen arguments. A web-fetch tool that returns a Markdown page can contain a hidden instruction that the model interprets as “call the database tool with these arguments.”
Mitigations:
- Schema validation on tool arguments: every tool call must be validated against its declared schema before execution. Free-form argument strings are a vulnerability.
- Authorization checks: tool calls must be authorized against the calling user’s permissions, not just the model’s. A model shouldn’t be able to bypass user-level access controls.
- Idempotency and dry-run: side-effecting tool calls should be idempotent or support dry-run mode, so an attacker-triggered unwanted call doesn’t cause damage.
- Rate limiting: per-tool-call rate limits prevent runaway execution.
The OpenTelemetry connection
The unifying layer across all these tools is OpenTelemetry. The OTel GenAI semantic conventions (still being finalized as of mid-2026) standardize how LLM and tool calls are represented in spans. Once your application emits OTel-conformant spans, any tracing platform that ingests OTel can consume them.
Practical implications:
- Lock-in is low: if you instrument with OTel, switching from Langfuse to Phoenix or to a hosted vendor is a config change, not a code change.
- Multi-vendor is feasible: run Langfuse for self-hosted traces, send the same OTel stream to a hosted vendor for evaluation, ship a subset to an observability SaaS like Datadog or Honeycomb for traditional APM.
- Custom instrumentation is straightforward: any code path that matters for debugging can be wrapped with
tracer.start_as_current_span(...)and the span will appear in the trace explorer.
For a new project in 2026, the right default is OTel-native instrumentation. The tracing platform is a deployment choice; the instrumentation is portable.
What to actually deploy
For a small team (1–5 engineers) shipping an agent pipeline:
- Tracing: Langfuse self-hosted, OpenTelemetry-instrumented. Single Docker Compose stack. Covers the “what is my agent doing in production” question.
- Evaluation: Opik or Langfuse’s built-in evaluation. Start with heuristic evals (response length, refusal rate, JSON validity), add LLM-as-judge for quality signals.
- Security: Secret redaction in the tracing platform, schema validation on all tool calls, egress allowlisting on the agent runtime.
For a mid-size team (5–50 engineers) with multiple agent products:
- Tracing: Langfuse or Phoenix self-hosted, with per-team project separation.
- Evaluation: Opik for evaluation depth, Langfuse for trace-side eval.
- Security: Add a prompt-injection test suite (Garak, PromptArmor, or custom) running in CI against every prompt and tool schema change.
For a large team (50+) with regulatory exposure:
- Tracing: Phoenix for its Elastic License and drift detection, plus Datadog/Honeycomb for traditional APM.
- Evaluation: Dedicated eval platform (Braintrust, Confident AI, or in-house).
- Security: Dedicated AI security tooling (PromptArmor, Lakera), prompt-injection red team, formal threat model for every new agent deployment.
Where to dig further
- Comet: AI Observability Tools for Agentic Systems in 2026 — clean taxonomy of the open-source tier (Opik, Langfuse, Phoenix, MLflow)
- Arize: 14 Best AI Observability Tools for Agents in 2026 — broader comparison including hosted vendors
- morphllm: AI Agent Observability Tools (2026) — focused on OpenTelemetry-native tools
- firecrawl: Best LLM Observability Tools in 2026 — Phoenix and Arize coverage
- marktechpost: Top LLM Observability and Evaluation Platforms in 2026 — Langfuse, LangSmith, Braintrust, Arize comparison
- The OpenTelemetry GenAI semantic conventions registry — the underlying standard that makes the tracing layer portable
The observability and security stack for AI agents has matured from “log and pray” into a real engineering category. The tools are real, the standards are converging on OpenTelemetry, and the threat model has a clear shape. If you’re building agent pipelines in 2026, this is the layer to invest in alongside the model selection.
Comments
Powered by GitHub Discussions via Giscus. Sign in with GitHub to leave a comment.