Skip to content

TechnicalBeginner

3D diagram: AI in cross-section

For: engineers and SREs · architects · team managers · business and productPrerequisites: None.

Reading mode

A generative AI application looks like a five-tier cake: data at the bottom, then compute infrastructure, the model, orchestration (prompts, document retrieval, agents) and finally the application the user sees. Each layer can drift, each layer can be attacked. This exploded view opens the stack: click a slab to read its card, switch between the signals to Observe and the Risks, then send a request to follow its trace from top to bottom and back.

Drag the stack to rotate it, then click a layer (or press Enter) to open its card. Layers go from data, at the bottom, to the application, at the top.

Data
Infrastructure
Model
Orchestration
Application

What to observe, layer by layer

What can go wrong, layer by layer

  • Application
    • User feedback and drop-offs
    • Content policy violations
    • Personal data detected in outputs
    • OWASP LLM02Sensitive information disclosure in outputs
    • Guardrail bypass
    • Loss of trust before any complaint
  • Orchestration
    • End-to-end trace of every request (OpenTelemetry)
    • Relevance of retrieved documents
    • Tokens and tool calls per request
    • OWASP LLM01Prompt injection
    • Pipeline drift: chunking or template changed silently
    • OWASP LLM06Agent with excessive agency
  • Model
    • Quality score from automated evaluation
    • Gap against a golden dataset
    • Exact model version on every call
    • OWASP LLM09Hallucinations and misinformation
    • Concept drift: the world changes, the model does not
    • Silent update at the provider
  • Infrastructure
    • Time to first token (TTFT)
    • Tokens per second, queue, GPU memory
    • Cost per request
    • Saturation and queues
    • OWASP LLM10Unbounded consumption: runaway cost
    • A busy GPU says nothing about quality
  • Data
    • Freshness and coverage of sources
    • Input distribution against the baseline
    • Personal data detected in logs
    • Data drift: user questions change
    • OWASP LLM04Data and model poisoning
    • Stale or incomplete sources

  1. Open the Model layer. It predicts the most likely next word without checking facts: its quality shows neither in latency nor in GPU utilization, but in an automated evaluation score.
  2. Switch to Risks mode and spot the OWASP IDs. Prompt injection (LLM01) targets orchestration, sensitive information disclosure (LLM02) the application, poisoning (LLM04) the data.
  3. Send a request. It crosses all five layers on the way down and back up: the trace is what ties their signals together and tells you where to look.

The same cross-section in 2 min 21, layer by layer, with a request crossing the whole stack.

Read the video text

Your dashboard is green. Yet your AI assistant has been answering wrong for weeks. An AI does not crash: it degrades silently. A generative AI application is a stack of five layers. Each one can drift, each one can be attacked. Let us open it, layer by layer, starting from the bottom.

Layer 1: Data. The fuel: internal documents, knowledge bases, training and RAG data.

What can go wrong: Data drift: user questions change ; Data and model poisoning (OWASP LLM04) ; Stale or incomplete sources.

What to observe: Freshness and coverage of sources ; Input distribution against the baseline ; Personal data detected in logs.

Layer 2: Infrastructure. The compute: GPUs, inference servers, cloud or cluster.

What can go wrong: Saturation and queues ; Unbounded consumption: runaway cost (OWASP LLM10) ; A busy GPU says nothing about quality.

What to observe: Time to first token (TTFT) ; Tokens per second, queue, GPU memory ; Cost per request.

Layer 3: Model. It predicts the most likely next word; it does not check facts.

What can go wrong: Hallucinations and misinformation (OWASP LLM09) ; Concept drift: the world changes, the model does not ; Silent update at the provider.

What to observe: Quality score from automated evaluation ; Gap against a golden dataset ; Exact model version on every call.

Layer 4: Orchestration. The glue: prompts, document retrieval (RAG), agents and tool calls.

What can go wrong: Prompt injection (OWASP LLM01) ; Pipeline drift: chunking or template changed silently ; Agent with excessive agency (OWASP LLM06).

What to observe: End-to-end trace of every request (OpenTelemetry) ; Relevance of retrieved documents ; Tokens and tool calls per request.

Layer 5: Application. What the user sees: assistant, chatbot, business feature.

What can go wrong: Sensitive information disclosure in outputs (OWASP LLM02) ; Guardrail bypass ; Loss of trust before any complaint.

What to observe: User feedback and drop-offs ; Content policy violations ; Personal data detected in outputs.

A request crosses all five layers, there and back. The trace ties their signals together: it tells you where to look.

Observing AI, from data to answer: www.meantimetolearn.com.

Music: “Radar Focus”, Blue Saga (Epidemic Sound).

Once the five layers are in place, the next step is to understand what to observe at each one: the guide Understanding generative AI observability sets out the signals, the system types and the maturity levels.

Sources: OWASP Top 10 for LLM Applications 2025.

Revised on 4 October 2026: progression box added, next step toward the guide.