Datadog Agent Observability
Datadog
Datadog module for end-to-end tracing of LLM and agent apps, with cost/latency dashboards, quality clustering and sensitive-data redaction.
Tracing, evaluation, prompt management, and cost monitoring for LLM apps and agents.
Observability tools record what models and agents actually did — traces, tool calls, cost, latency — and evaluation tools score it. Compare self-hosted and SaaS options on license terms and data residency.
Tools from the major platform vendors are listed first; independent and open-source options follow below.
Datadog
Datadog module for end-to-end tracing of LLM and agent apps, with cost/latency dashboards, quality clustering and sensitive-data redaction.
6 matching non-vendor tools in this filtered view.
Arize AI
Source-available AI observability and evaluation platform: OpenTelemetry/OpenInference tracing, evals, prompt iteration and experiments.
Braintrust
Commercial AI observability and evaluation platform for tracing agents, running evals and experiments, and building datasets from production logs.
Langfuse (ClickHouse)
Open-source LLM engineering platform for tracing, evaluations, prompt management and cost/latency dashboards, built on OpenTelemetry.
LangChain
Framework-agnostic platform from LangChain for tracing, monitoring, evaluating and deploying LLM apps and agents, with OpenTelemetry ingest.
Linux Foundation (MLflow project)
Open-source AI engineering platform with OpenTelemetry-compatible GenAI tracing, LLM-judge evaluation, prompt registry and AI gateway.
Comet
Open-source platform to trace, evaluate and monitor LLM apps, RAG systems and agents, with LLM-as-a-judge metrics and prompt optimization.
MLflow 3.16.1 removed the default basic-auth admin password (a security fix for self-hosted servers) and added scorer timeouts.
SourceRecording what an LLM application or agent actually did — prompts, model calls, tool calls, retrieved context, latency, token cost, and errors — as traces you can search and replay. Most tools here use or accept OpenTelemetry so traces can flow into existing monitoring.
Observability shows what happened; evaluation scores whether it was good, using test datasets, automated judges, or human review, before and after release. Most tools here now do both, but they differ in whether evals or tracing is the core of the product.
Traces often contain full prompts, retrieved documents, and outputs, so they carry the same sensitivity as the underlying data. Self-hostable options keep that data in your environment; check each per-tool page for the license (several are open-core or source-available) and SaaS data-residency options.
Compare the active category against the platform foundation layer, the cross-category updates feed, and the sourcing/contribution guide.
Review Microsoft Foundry, Amazon Bedrock, and Gemini Enterprise Agent Platform as the foundation layer behind this category.
Check the market-intelligence feed for high-impact moves, plus the expandable full log for releases, deprecations, acquisitions, and other notable changes.
See sourcing standards, contribution rules, and project scope before adding or updating tracked tools.