AI Observability & Evaluation

Tracing, evaluation, prompt management, and cost monitoring for LLM apps and agents.

Observability tools record what models and agents actually did — traces, tool calls, cost, latency — and evaluation tools score it. Compare self-hosted and SaaS options on license terms and data residency.

In dataset
7
Filter tools by type
Filter tools by cloud support
7 matching tools
More filters

Major vendor tools

Tools from the major platform vendors are listed first; independent and open-source options follow below.

Filtered open source and third-party tools

6 matching non-vendor tools in this filtered view.

Open Source

Source-available AI observability and evaluation platform: OpenTelemetry/OpenInference tracing, evals, prompt iteration and experiments.

Elastic License 2.0
Self-hostedLicense high

Braintrust

Braintrust

Commercial

Commercial AI observability and evaluation platform for tracing agents, running evals and experiments, and building datasets from production logs.

Proprietary
SaaSHybridSOC 2License medium

Langfuse

Langfuse (ClickHouse)

Open Source

Open-source LLM engineering platform for tracing, evaluations, prompt management and cost/latency dashboards, built on OpenTelemetry.

MIT core + EE paths
SaaSSelf-hostedSOC 2License medium

LangSmith

LangChain

Commercial

Framework-agnostic platform from LangChain for tracing, monitoring, evaluating and deploying LLM apps and agents, with OpenTelemetry ingest.

Proprietary
SaaSHybridSelf-hostedLicense medium

MLflow

Linux Foundation (MLflow project)

Open Source

Open-source AI engineering platform with OpenTelemetry-compatible GenAI tracing, LLM-judge evaluation, prompt registry and AI gateway.

Apache 2.0
Self-hostedSaaSLicense low

Opik

Comet

Open Source

Open-source platform to trace, evaluate and monitor LLM apps, RAG systems and agents, with LLM-as-a-judge metrics and prompt optimization.

Apache 2.0
SaaSSelf-hostedSOC 2License low

Important notes

Warning

Arize Phoenix: Phoenix uses the Elastic License 2.0, which is source-available rather than OSI open source. It forbids offering Phoenix to third parties as a hosted or managed service and forbids circumventing license-key features. Internal self-hosting is allowed.

Warning

Datadog Agent Observability: Datadog's docs now title the product "Agent Observability" and still use "LLM Observability" interchangeably. Not available on the US1-FED/US2-FED government sites (checked 2026-09-30).

Warning

Langfuse: Langfuse is open core: everything outside the ee/, web/src/ee/ and worker/src/ee/ directories is MIT, but code in those directories needs a commercial Langfuse Enterprise License for production use (development and testing are allowed without one).
2026-09-17
MLflow
MLflow

MLflow 3.16.1 removed the default basic-auth admin password (a security fix for self-hosted servers) and added scorer timeouts.

Source

Frequently asked questions

  • What is AI observability?

    Recording what an LLM application or agent actually did — prompts, model calls, tool calls, retrieved context, latency, token cost, and errors — as traces you can search and replay. Most tools here use or accept OpenTelemetry so traces can flow into existing monitoring.

  • How does evaluation differ from observability?

    Observability shows what happened; evaluation scores whether it was good, using test datasets, automated judges, or human review, before and after release. Most tools here now do both, but they differ in whether evals or tracing is the core of the product.

  • Should traces be kept in-house?

    Traces often contain full prompts, retrieved documents, and outputs, so they carry the same sensitivity as the underlying data. Self-hostable options keep that data in your environment; check each per-tool page for the license (several are open-core or source-available) and SaaS data-residency options.

Explore adjacent hubs

Compare the active category against the platform foundation layer, the cross-category updates feed, and the sourcing/contribution guide.