Cross-cutting
Glossary
Each term is given with its French equivalent, for readers moving between both versions of the site.
ADKAR: change management model published by Prosci, five stages each person goes through: awareness, desire, knowledge, ability, reinforcement.
Burn rate (taux de consommation): how fast an SLO’s error budget is consumed. A burn rate of 1 exhausts the budget exactly at the end of the window.
Cardinality (cardinalité): number of distinct time series a metric produces, that is, unique combinations of its labels. A major cost driver of a metrics system.
Chargeback (refacturation interne): actually charging the cost of observability or cloud to each team’s budget.
Collector: component that receives, transforms and exports telemetry. Usually refers to the OpenTelemetry Collector.
Data contract (contrat de données): written rule, checked in continuous integration, defining for a signal its allowed attributes, cardinality budget, retention and destination.
DCGM (Data Center GPU Manager): NVIDIA GPU monitoring tool; its exporter exposes per-GPU utilization, memory, temperature and power in Prometheus format.
Drift (dérive): progressive degradation of a model in production. Data drift (questions change), concept drift (reality changes) and pipeline drift (a component changes).
eBPF: Linux kernel technology that runs verified programs on hooks (syscalls, functions, network) to observe or filter without changing code.
Error budget (budget d’erreur): the share of unreliability an SLO tolerates over a period. A 99.9% SLO over 30 days leaves about 43 minutes of downtime.
Essential entity (entité essentielle): category of organizations under NIS2 with the strictest regime, depending on sector and size; maximum fine of at least €10M or 2% of worldwide turnover.
Exemplar: link from a metric data point to a specific trace, to go from a latency spike to the request that caused it.
Faithfulness (fidélité): share of a response’s claims actually supported by the retrieved context. The core quality metric of a RAG system.
Forge: a platform hosting code repositories (Git), with issue and release tracking, such as GitLab or GitHub. On this site, “the forge” means the Labs & Trainings forge, where the lab code is freely available.
Golden set (jeu de référence): a human-annotated set of questions and answers, used for offline evaluation and to calibrate an LLM judge.
Immutable storage (stockage immuable): storage, also called WORM (write once, read many), where written data can no longer be changed or deleted before the end of its retention period; a necessary condition for evidence, not a sufficient one (see log integrity).
Important entity (entité importante): category of organizations under NIS2 with after-the-fact supervision; maximum fine of at least €7M or 1.4% of worldwide turnover.
Interconnect: very fast network linking the nodes of an HPC cluster, such as InfiniBand.
LLM as a judge (LLM juge): using a language model to score another model’s answers. Fast and reproducible, but subject to known biases.
Log (journal): timestamped record of a discrete event, structured or not.
Metric (métrique): numeric measurement aggregated over time, identified by a name and labels.
MIG (Multi-Instance GPU): partitioning of an NVIDIA GPU into instances isolated in compute and memory: each GPU instance has its own memory paths (L2 cache, controllers, bandwidth), according to the NVIDIA documentation. Compute instances created inside the same GPU instance share its memory and engines (MIG concepts).
MTTL (Mean Time To Learn): the average time it takes to understand what is going on, which comes before any fix. The name of this site, echoing MTTR.
MTTR (mean time to recovery): average time to restore service after an incident.
NCCL: NVIDIA collective communication library between GPUs (AllReduce, broadcast) used by distributed training.
OTLP (OpenTelemetry Protocol): the telemetry transport protocol defined by OpenTelemetry.
Payback period (délai de récupération): number of months needed for the cumulative gains of an investment to equal its cost.
Profile (profil): sampling of program activity (CPU, memory) to locate expensive code. The fourth OpenTelemetry signal, still in development (not stable) as of 2 October 2026 according to the specification status page.
RACI: matrix showing, for each activity, who does the work (R), who is accountable and signs off (A, only one per activity), who is consulted (C) and who is informed (I).
RAG (retrieval augmented generation): retrieving relevant documents and passing them to the model as context.
RDMA (Remote Direct Memory Access): direct access by the network adapter to another machine’s memory, bypassing the kernel; the basis of InfiniBand and RoCE.
SBOM (software bill of materials): the list of software components included in a product, with their versions.
Semantic convention (convention sémantique): standardized names and meanings of OpenTelemetry attributes, such as http.response.status_code or gen_ai.operation.name.
Showback (affichage des coûts): showing each team what it consumes, without charging its budget; the first step towards accountability.
Silent failure (défaillance silencieuse): a failure with no technical error (hallucination, cost drift, agent loop) that accumulates until it is detected late.
SLI (service level indicator): a measure of service quality as seen by the user, for example the share of requests served under 300 ms.
SLO (service level objective): a target set on an SLI, for example 99.5% of requests under 300 ms over 30 days.
Slurm: the reference HPC scheduler; it queues jobs, allocates resources and accounts for usage per account.
Span: a unit of work in a trace, with a start, an end, attributes and a parent.
Straggler (traînard): a GPU or node slower than the others, which makes the whole cluster wait at every synchronization barrier.
Tail sampling (échantillonnage en queue): the decision to keep a trace or not, made once the trace is complete. Keeps every error and a fraction of nominal traffic.
Tool poisoning (empoisonnement d’outil): malicious instructions hidden in an MCP tool definition to manipulate the agent that uses it.
Trace: the set of spans describing a request’s path through a distributed system.
TTFT (time to first token): time between sending a request and receiving the first generated token. The perceived-latency indicator for streaming.
Wilson bound (borne de Wilson): a confidence interval bound for a proportion that holds on small samples. Used to compare a quality SLO with its target without false precision.
Revised on 2 October 2026: cardinality described as a major cost driver rather than the first one, MIG definition clarified and sourced from the NVIDIA documentation.