Technical
Part III. Measurement grid per component
For: engineers and SREs · architectsPrerequisites: Have read part I of the method.
For each component: a diagram, then the systematic grid what / why / impact if unmonitored. The (S) marker flags a silent failure.
8. LLM (the inference brick)
Section titled “8. LLM (the inference brick)”flowchart LR REQ["Request"] --> QUEUE["Queue"] --> PREFILL["Prefill"] --> DECODE["Decode (streaming)"] --> RESP["Response"] GPU["GPU / KV cache / batch"] -.-> PREFILL GPU -.-> DECODE
| What | Why (decision) | Impact if unmonitored |
|---|---|---|
| TTFT, E2E latency (p95, p99) | Hold the perceived latency budget | Degradation unnoticed until complaints |
| Throughput tokens/s, requests/s | Size, autoscale | Unanticipated saturation, capacity incident |
| GPU saturation, KV cache, batch (DCGM, vLLM) | Anticipate degradation | (S) Eviction and queue rise silently, then an unexplained spike |
Error ratio by error.type | Isolate timeout, rate-limit, content_filter | Blind diagnosis, longer resolution |
finish_reasons distribution | Detect truncated or blocked answers | (S) Cut answers served with no alert |
| Refusal rate | Detect a behavior shift | (S) The product refuses more and more, no signal |
| Hallucination, faithfulness (sampled eval) | Guarantee correctness | (S) The system asserts confidently false claims, business and legal risk |
| Output drift | Detect slow evolution | (S) Invisible regression over weeks |
| Tokens in/out, cost per request | Govern the budget | (S) Bill discovered at month end |
| Cache hit rate | Measure a direct cost lever | Permanent overspend unidentified |
| Prompt bloat (input tokens over time) | Detect context growth | (S) Cost and latency drift slowly |
9. RAG
Section titled “9. RAG”flowchart LR ING["Ingestion"] --> CHUNK["Chunking"] --> EMB["Embedding"] --> IDX["Vector index"] Q["Question"] --> RET["Retrieval"] --> RANK["Reranking"] --> CTX["Context"] --> GEN["Generation"] --> ANS["Answer"] IDX --> RET
| What | Why (decision) | Impact if unmonitored |
|---|---|---|
| Context precision | Are the retrieved chunks relevant | (S) Off-topic context, hallucination downstream |
| Context recall | Did you retrieve everything needed | (S) Incomplete or invented answers |
| Hit rate, MRR | Is the right document retrieved and well ranked | You fix generation by mistake |
| Faithfulness / groundedness | Answer grounded in context | (S) Hallucinations served as facts |
| Answer relevance | Does it answer the question | Off-target answers unnoticed |
| Citation accuracy | Cited sources correct | (S) False references, compliance risk |
| Coverage (questions with no context) | Measure knowledge base gaps | (S) A documentary blind spot that widens |
| Base freshness | Knowledge not stale | (S) Outdated answers served confidently |
| Retrieval latency, index health | Hold the chain budget | Degradation mis-attributed |
| Embeddings and vector store cost | Govern cost | Uncontrolled infrastructure overspend |
Eval frameworks: RAGAS and DeepEval for gating, Arize Phoenix for drift and embedding cluster visualization. The evaluator problem (non-deterministic judge, drift, calibration) is covered in section 15.
10. Agent
Section titled “10. Agent”flowchart TB
START["invoke_agent"] --> PLAN["Planning"] --> SEL["Tool selection"] --> TOOL["execute_tool"] --> OBS["Observation"]
OBS --> DEC{"Task done?"}
DEC -->|"no"| SEL
DEC -->|"yes"| END["Final answer"]
DEC -.->|"risk: loop"| LOOP["Loop detection"]
| What | Why (decision) | Impact if unmonitored |
|---|---|---|
| Task success rate | Measure real end-to-end work | (S) Partial failure with no error, product value eroded |
| Steps per task | Detect abnormal trajectories | (S) Cost and latency drift from lengthening |
| Loop detection | Spot repetition | (S) A loop that burns the budget, seen on the bill |
| Tool failure rate | Identify failing tools | Costly retries, root cause not isolated |
| Latency per step and per task | Decompose the end-to-end path | Diagnosis impossible on long trajectories |
| Cost per task (cumulative tokens) | Know the real cost of a task | (S) Mispriced offering |
| Tool selection accuracy | The right tool at the right time | (S) Degraded results with no signal |
| Human handoff rate | Measure real autonomy | Operator load underestimated, ROI distorted |
| Guardrails, injection attempts | Agentic security | Command injection undetected |
Tools: Laminar and Langfuse (transcript, debug), Phoenix and Confident AI (multi-step eval), LangSmith and LangGraph Studio for a LangGraph stack.
11. MCP
Section titled “11. MCP”flowchart TB P["Protocol layer: JSON-RPC over transport, sessions, schemas"] --> T["Execution layer: tool calls, latency, errors"] --> A["Task layer: impact on agent success"] A -.->|"cascade"| P
| What | Why (decision) | Impact if unmonitored |
|---|---|---|
| Latency and errors per tool and server | Execution health, isolation | Failure mis-attributed to the agent |
| Session health, client-server desync | Transport stability | (S) Desync that degrades the agent with no error |
| Authentication failures per identity | Detect abuse and misconfiguration | Illegitimate access, secret exposure |
| Enumeration / discovery spikes | Detect hostile reconnaissance | (S) Invisible reconnaissance before exploitation |
| Server mutations | Detect rug pull and schema poisoning | (S) A trusted tool turned malicious |
| Tool pinning by hash | Tool integrity over time | (S) Undetected modification |
| Invocation log | Regulatory audit trail | No traceability, non-compliance |
| Availability per server, health checks | Guarantee fallback | Agent looping on an unavailable server |
Security: the OWASP MCP Top 10 reference (2025 edition, in beta) defines ten categories: MCP01 token mismanagement and secret exposure, MCP02 privilege escalation via scope creep, MCP03 tool poisoning, MCP04 software supply chain attacks and dependency tampering, MCP05 command injection and execution, MCP06 prompt injection via contextual payloads, MCP07 insufficient authentication and authorization, MCP08 lack of audit and telemetry, MCP09 shadow MCP servers, MCP10 context injection and over-sharing. Rug pull, tool shadowing and schema poisoning are attack techniques described in the literature, mostly falling under MCP03; they are not categories of the reference. The grid above directly addresses MCP08.
Tool call auto-approval removes the last barrier that does not depend on the model. The MCPTox benchmark (2025) ran 20 LLM agents against 1,348 tool poisoning cases built on 45 real MCP servers: the attack succeeds in 36.5% of cases on average, up to 72.8% for the most exposed model, and no model refuses more than 3% of attacks. The benchmark measures model susceptibility, not the effect of human approval; I still recommend human-in-the-loop validation on sensitive tools, because it does not rely on the judgment of the model under attack. Tools: mcp-scan (Invariant Labs / Snyk) and Cisco mcp-scanner (YARA), complementary, plus an MCP gateway with audit and OTel export.
Revised on 2 October 2026: official OWASP MCP Top 10 categories (2025 edition, in beta) and MCPTox benchmark figures in place of an unsourced claim about auto-approval.