Technical
6. The tools landscape
For: engineers and SREs · architectsPrerequisites: Have read chapter 4 of the guide.
The AI observability ecosystem is dense and moving fast. The table below organizes widely used options by category and hosting mode, in alphabetical order. Inclusion does not imply endorsement: selection depends on the specific constraints of the deployment. This chapter is the site’s reference for the tools landscape and for licenses (section 6.4); other pieces link here.
| Category | Self-hostable software | Managed services (SaaS) |
|---|---|---|
| Instrumentation | OpenInference (Arize), OpenLLMetry (Traceloop), OpenTelemetry GenAI SDKs | Vendor SDKs (Arize, Datadog, New Relic) |
| Trace backend | Grafana Tempo, Jaeger, Langfuse, OpenObserve, Phoenix, SigNoz | Arize AX, Datadog LLM Observability, Langfuse Cloud, New Relic AI Monitoring |
| Metrics backend | Grafana Mimir, Prometheus, Thanos, VictoriaMetrics | Chronosphere, Datadog, Grafana Cloud, New Relic |
| Logs backend | Elasticsearch, Grafana Loki, OpenObserve, VictoriaLogs | Datadog, Elastic Cloud, Grafana Cloud, Splunk |
| Evaluation framework | DeepEval, OpenAI Evals, Phoenix evals, promptfoo, Ragas | Arize, Braintrust, Galileo, Patronus AI |
| Drift and embeddings | Evidently, NannyML, Phoenix, WhyLabs | Arize, Fiddler |
| Visualization | Apache Superset, Grafana, Perses, Phoenix UI | Datadog dashboards, Grafana Cloud, New Relic dashboards |
| Prompt management | Langfuse, Phoenix | Langfuse Cloud, PromptLayer, Vellum |
6.1 Choosing between self-hosting and a managed service
Section titled “6.1 Choosing between self-hosting and a managed service”Three factors dominate the decision.
Data sovereignty
Section titled “Data sovereignty”Regulated sectors (banking, defense, healthcare, public sector) often require on-premise or sovereign-cloud deployment. Neither NIS2 nor DORA by itself requires the observability backend to reside in a given jurisdiction. That requirement can, however, come from an internal policy, a sectoral regulator or a contract. DORA requires financial entities to manage ICT third-party risk and to keep a register of information that records, among other things, the countries where data is stored and processed (Implementing Regulation (EU) 2024/2956): an observability backend that stores prompts belongs in it like any other service. When a location constraint applies, the field narrows to self-hosted software or vendors with compliant regions, and audit trails must cover access to stored prompts.
Commercial LLM observability is priced per span, per token or per user. At meaningful production volume the observability cost can rival the LLM cost itself. A rough sizing: 100 spans per request at 1,000 RPS sustained gives 8.64 billion spans per day, or about 3,150 billion spans per year, for tracing alone. Multiply this volume by your provider’s price per span (or per million spans). Prices change fast and vary by contract: this guide gives none; use your provider’s dated price list.
Self-hosting is not free either: it shifts the cost to infrastructure and operating time, which deserve the same scrutiny.
Integration
Section titled “Integration”The existing observability stack often dictates the AI observability choice. A site already running Grafana, Prometheus and Tempo can keep its backends and add GenAI instrumentation plus an AI-specific exploration tool (Phoenix, Langfuse or equivalent). A site already on Datadog or New Relic can enable its vendor’s LLM module. Every additional tool adds operational load (software to maintain, a data model, access rights): add one only when it answers a question the current stack cannot, and keep the OpenTelemetry Collector as the common entry point.
6.2 A self-hosted reference stack
Section titled “6.2 A self-hosted reference stack”This is the stack I recommend as a starting point for European deployments with sovereignty requirements. It is a proposal to adapt: each layer is replaceable, and other combinations meet the same criteria.

The figure, from the first version of this guide, is titled “open source stack”: Phoenix, which appears in it, is under a source-available license (Elastic License 2.0), see above.
Figure 9. Reference architecture for a self-hosted stack. Each layer is replaceable independently. The OpenTelemetry Collector is the integration point: it decouples instrumentation from backends.
Component selection criteria
Section titled “Component selection criteria”- Metrics (VictoriaMetrics in the figure): worth considering if your AI metrics have high cardinality and you want a single binary with long retention. The vendor states better cardinality handling than Prometheus; measure it on your own workload. Alternatives: Prometheus with remote storage, Thanos or Grafana Mimir.
- AI traces and evaluation (Phoenix): AI-specific trace and evaluation UI in one tool, which reduces the integration surface. Source-available license (ELv2), to be validated. Open source alternative: Langfuse (MIT for the core).
- General-purpose traces (Tempo): distributed tracing for non-AI services, useful when the platform mixes traditional services and AI features. Alternatives: Jaeger, SigNoz.
- Logs (VictoriaLogs): single binary, operated much like VictoriaMetrics, but with its own query language (LogsQL), distinct from MetricsQL. Alternatives: Grafana Loki, OpenObserve, Elasticsearch.
- Instrumentation (OpenInference in the figure): a trade-off, not an unconditional recommendation. OpenInference covers many LLM frameworks, exports over OTLP and is Phoenix’s native convention. But it uses its own attributes (
openinference.span.kind,llm.*) rather than the OpenTelemetrygen_ai.*attributes this guide takes as its reference. Two consistent paths: OpenInference if Phoenix is your main exploration tool and you accept its convention, converting togen_ai.*with the Collector’stransformprocessor where other backends need it; or the native OpenTelemetry GenAI instrumentations (or OpenLLMetry), which emitgen_ai.*, if you favor a single convention portable across tools, at the price of conventions still in development. Either way, a single convention should feed dashboards and alerts. - RAG evaluation (Ragas): Python framework dedicated to RAG metrics. Compare with DeepEval or Phoenix evals depending on the metrics you need.
- Visualization (Grafana): the operational visualization layer. The AI trace tool’s UI handles AI-specific exploration.
6.3 A managed-service stack
Section titled “6.3 A managed-service stack”For organizations without location constraints and with an existing commercial observability vendor, four common configurations.
- Already on a general-purpose vendor (Datadog, New Relic or equivalent): enable its LLM observability module. Single billing, single UI; check evaluator coverage and cost at volume.
- Quality and evaluation are the priority: a specialized platform such as Arize AX or Braintrust, or Langfuse Cloud. Compare them on built-in evaluators, the ability to write your own, and data export.
- Evaluation specifically is the priority, tracing secondary: Galileo, Patronus AI or Braintrust.
- Prompt management is the central need: Langfuse Cloud, PromptLayer or Vellum.
6.4 Component licenses
Section titled “6.4 Component licenses”This section is the site’s reference on licenses; other pieces link here. “Open source” covers very different regimes, and some so-called “source-available” licenses are not part of it. The confusion is paid for at industrialization time. Licenses change: check the one for the version you deploy.
| Component | Role | License | Point of attention |
|---|---|---|---|
| OpenTelemetry Collector | gateway | Apache-2.0 | none |
| Jaeger | traces | Apache-2.0 | none |
| Prometheus | metrics | Apache-2.0 | none |
| VictoriaMetrics | metrics | Apache-2.0 (community edition, cluster version included) | downsampling, multiple retentions and anomaly detection reserved for the enterprise edition |
| ClickHouse | columnar storage | Apache-2.0 | none |
| OpenLIT, OpenLLMetry | instrumentation | Apache-2.0 | none |
| Ragas, DeepEval | evaluation | Apache-2.0 | cost of judge calls |
| Grafana, Tempo, Loki | visualization, traces, logs | AGPL-3.0 | obligations on redistribution and also when a modified version is made available to users over the network |
| Langfuse | LLM observability | MIT for the core | some features under a commercial license; acquired by ClickHouse on 16 January 2026 |
| Arize Phoenix | LLM observability | Elastic License 2.0 (source-available) | not open source in the OSI sense: offering the software as a managed service to third parties is forbidden |
| Elasticsearch | logs | source code of the free features under a choice of AGPL v3, SSPL or ELv2 | official binary distributions under ELv2 |
Two regimes cause trouble in legal review. The Elastic License 2.0 forbids offering the product as a managed service to third parties: no effect for internal use, a blocker for a vendor that embeds the component in its offering. The AGPL of Grafana, Tempo and Loki is generally not an obstacle for an unmodified internal deployment; it does apply to redistribution and to network interaction with a modified version: modifying Grafana and opening it to users, even without distributing it, obliges you to offer them the modified source code. Flag it before someone builds a product on top of it.
To choose the LLM-specific component, apply three criteria the same way to every option: the license (open source in the OSI sense, or source-available), the ability to self-host, and the ability to receive OTLP with the gen_ai.* conventions without a proprietary SDK. The production article applies these criteria to its own stack.
Next: 7. Operating the platform.
Revised on 2 October 2026: prices removed, Humanloop removed (shut down), Langfuse and Traceloop acquisitions, WhyLabs now open source, Phoenix and Elasticsearch licenses clarified, neutral table columns, NIS2 and DORA location requirements corrected.
Revised on 4 October 2026: chapter designated as the site’s reference for tools and licenses, section 6.4 on per-component licenses (taken over from the production article), OpenInference choice presented as a trade-off against the gen_ai.* conventions.