For: engineers and SREs · architects · team managers · business and productPrerequisites: None.
Reading mode
Decision-maker reading: bridges first.
A generative AI application looks like a five-tier cake: data at the bottom, then compute infrastructure, the model, orchestration (prompts, document retrieval, agents) and finally the application the user sees. Each layer can drift, each layer can be attacked. This exploded view opens the stack: click a slab to read its card, switch between the signals to Observe and the Risks, then send a request to follow its trace from top to bottom and back.
Drag the stack to rotate it, then click a layer (or press Enter) to open its card. Layers go from data, at the bottom, to the application, at the top.
Data
Freshness and coverage of sources
Input distribution against the baseline
Personal data detected in logs
Data drift: user questions change
LLM04Data and model poisoning
Stale or incomplete sources
Infrastructure
Time to first token (TTFT)
Tokens per second, queue, GPU memory
Cost per request
Saturation and queues
LLM10Unbounded consumption: runaway cost
A busy GPU says nothing about quality
Model
Quality score from automated evaluation
Gap against a golden dataset
Exact model version on every call
LLM09Hallucinations and misinformation
Concept drift: the world changes, the model does not
Silent update at the provider
Orchestration
End-to-end trace of every request (OpenTelemetry)
Relevance of retrieved documents
Tokens and tool calls per request
LLM01Prompt injection
Pipeline drift: chunking or template changed silently
LLM06Agent with excessive agency
Application
User feedback and drop-offs
Content policy violations
Personal data detected in outputs
LLM02Sensitive information disclosure in outputs
Guardrail bypass
Loss of trust before any complaint
What to observe, layer by layer
What can go wrong, layer by layer
Application
User feedback and drop-offs
Content policy violations
Personal data detected in outputs
OWASP LLM02Sensitive information disclosure in outputs
Guardrail bypass
Loss of trust before any complaint
Orchestration
End-to-end trace of every request (OpenTelemetry)
Relevance of retrieved documents
Tokens and tool calls per request
OWASP LLM01Prompt injection
Pipeline drift: chunking or template changed silently
OWASP LLM06Agent with excessive agency
Model
Quality score from automated evaluation
Gap against a golden dataset
Exact model version on every call
OWASP LLM09Hallucinations and misinformation
Concept drift: the world changes, the model does not
Silent update at the provider
Infrastructure
Time to first token (TTFT)
Tokens per second, queue, GPU memory
Cost per request
Saturation and queues
OWASP LLM10Unbounded consumption: runaway cost
A busy GPU says nothing about quality
Data
Freshness and coverage of sources
Input distribution against the baseline
Personal data detected in logs
Data drift: user questions change
OWASP LLM04Data and model poisoning
Stale or incomplete sources
Layer 1 of 5
Data
What it is
The fuel: internal documents, knowledge bases, training and RAG data.
What can go wrong
Data drift: user questions change
Data and model poisoning OWASP LLM04
Stale or incomplete sources
What to observe
Freshness and coverage of sources
Input distribution against the baseline
Personal data detected in logs
Layer 2 of 5
Infrastructure
What it is
The compute: GPUs, inference servers, cloud or cluster.
What can go wrong
Saturation and queues
Unbounded consumption: runaway cost OWASP LLM10
A busy GPU says nothing about quality
What to observe
Time to first token (TTFT)
Tokens per second, queue, GPU memory
Cost per request
Layer 3 of 5
Model
What it is
It predicts the most likely next word; it does not check facts.
What can go wrong
Hallucinations and misinformation OWASP LLM09
Concept drift: the world changes, the model does not
Silent update at the provider
What to observe
Quality score from automated evaluation
Gap against a golden dataset
Exact model version on every call
Layer 4 of 5
Orchestration
What it is
The glue: prompts, document retrieval (RAG), agents and tool calls.
What can go wrong
Prompt injection OWASP LLM01
Pipeline drift: chunking or template changed silently
Agent with excessive agency OWASP LLM06
What to observe
End-to-end trace of every request (OpenTelemetry)
Relevance of retrieved documents
Tokens and tool calls per request
Layer 5 of 5
Application
What it is
What the user sees: assistant, chatbot, business feature.
What can go wrong
Sensitive information disclosure in outputs OWASP LLM02
Open the Model layer. It predicts the most likely next word without checking facts: its quality shows neither in latency nor in GPU utilization, but in an automated evaluation score.
Switch to Risks mode and spot the OWASP IDs. Prompt injection (LLM01) targets orchestration, sensitive information disclosure (LLM02) the application, poisoning (LLM04) the data.
Send a request. It crosses all five layers on the way down and back up: the trace is what ties their signals together and tells you where to look.
The same cross-section in 2 min 21, layer by layer, with a request crossing the whole stack.
Read the video text
Your dashboard is green. Yet your AI assistant has been answering wrong for weeks. An AI does not crash: it degrades silently. A generative AI application is a stack of five layers. Each one can drift, each one can be attacked. Let us open it, layer by layer, starting from the bottom.
Layer 1: Data. The fuel: internal documents, knowledge bases, training and RAG data.
What can go wrong: Data drift: user questions change ; Data and model poisoning (OWASP LLM04) ; Stale or incomplete sources.
What to observe: Freshness and coverage of sources ; Input distribution against the baseline ; Personal data detected in logs.
Layer 2: Infrastructure. The compute: GPUs, inference servers, cloud or cluster.
What can go wrong: Saturation and queues ; Unbounded consumption: runaway cost (OWASP LLM10) ; A busy GPU says nothing about quality.
What to observe: Time to first token (TTFT) ; Tokens per second, queue, GPU memory ; Cost per request.
Layer 3: Model. It predicts the most likely next word; it does not check facts.
What can go wrong: Hallucinations and misinformation (OWASP LLM09) ; Concept drift: the world changes, the model does not ; Silent update at the provider.
What to observe: Quality score from automated evaluation ; Gap against a golden dataset ; Exact model version on every call.
Layer 4: Orchestration. The glue: prompts, document retrieval (RAG), agents and tool calls.
What can go wrong: Prompt injection (OWASP LLM01) ; Pipeline drift: chunking or template changed silently ; Agent with excessive agency (OWASP LLM06).
What to observe: End-to-end trace of every request (OpenTelemetry) ; Relevance of retrieved documents ; Tokens and tool calls per request.
Layer 5: Application. What the user sees: assistant, chatbot, business feature.
What can go wrong: Sensitive information disclosure in outputs (OWASP LLM02) ; Guardrail bypass ; Loss of trust before any complaint.
What to observe: User feedback and drop-offs ; Content policy violations ; Personal data detected in outputs.
A request crosses all five layers, there and back. The trace ties their signals together: it tells you where to look.
Observing AI, from data to answer: www.meantimetolearn.com.
Once the five layers are in place, the next step is to understand what to observe at each one: the guide Understanding generative AI observability sets out the signals, the system types and the maturity levels.