Skip to content

TechnicalExpert

Part III. Measurement grid per component

For: engineers and SREs · architectsPrerequisites: Have read part I of the method.

For each component: a diagram, then the systematic grid what / why / impact if unmonitored. The (S) marker flags a silent failure.

flowchart LR
  REQ["Request"] --> QUEUE["Queue"] --> PREFILL["Prefill"] --> DECODE["Decode (streaming)"] --> RESP["Response"]
  GPU["GPU / KV cache / batch"] -.-> PREFILL
  GPU -.-> DECODE
WhatWhy (decision)Impact if unmonitored
TTFT, E2E latency (p95, p99)Hold the perceived latency budgetDegradation unnoticed until complaints
Throughput tokens/s, requests/sSize, autoscaleUnanticipated saturation, capacity incident
GPU saturation, KV cache, batch (DCGM, vLLM)Anticipate degradation(S) Eviction and queue rise silently, then an unexplained spike
Error ratio by error.typeIsolate timeout, rate-limit, content_filterBlind diagnosis, longer resolution
finish_reasons distributionDetect truncated or blocked answers(S) Cut answers served with no alert
Refusal rateDetect a behavior shift(S) The product refuses more and more, no signal
Hallucination, faithfulness (sampled eval)Guarantee correctness(S) The system asserts confidently false claims, business and legal risk
Output driftDetect slow evolution(S) Invisible regression over weeks
Tokens in/out, cost per requestGovern the budget(S) Bill discovered at month end
Cache hit rateMeasure a direct cost leverPermanent overspend unidentified
Prompt bloat (input tokens over time)Detect context growth(S) Cost and latency drift slowly
flowchart LR
  ING["Ingestion"] --> CHUNK["Chunking"] --> EMB["Embedding"] --> IDX["Vector index"]
  Q["Question"] --> RET["Retrieval"] --> RANK["Reranking"] --> CTX["Context"] --> GEN["Generation"] --> ANS["Answer"]
  IDX --> RET
WhatWhy (decision)Impact if unmonitored
Context precisionAre the retrieved chunks relevant(S) Off-topic context, hallucination downstream
Context recallDid you retrieve everything needed(S) Incomplete or invented answers
Hit rate, MRRIs the right document retrieved and well rankedYou fix generation by mistake
Faithfulness / groundednessAnswer grounded in context(S) Hallucinations served as facts
Answer relevanceDoes it answer the questionOff-target answers unnoticed
Citation accuracyCited sources correct(S) False references, compliance risk
Coverage (questions with no context)Measure knowledge base gaps(S) A documentary blind spot that widens
Base freshnessKnowledge not stale(S) Outdated answers served confidently
Retrieval latency, index healthHold the chain budgetDegradation mis-attributed
Embeddings and vector store costGovern costUncontrolled infrastructure overspend

Eval frameworks: RAGAS and DeepEval for gating, Arize Phoenix for drift and embedding cluster visualization. The evaluator problem (non-deterministic judge, drift, calibration) is covered in section 15.

flowchart TB
  START["invoke_agent"] --> PLAN["Planning"] --> SEL["Tool selection"] --> TOOL["execute_tool"] --> OBS["Observation"]
  OBS --> DEC{"Task done?"}
  DEC -->|"no"| SEL
  DEC -->|"yes"| END["Final answer"]
  DEC -.->|"risk: loop"| LOOP["Loop detection"]
WhatWhy (decision)Impact if unmonitored
Task success rateMeasure real end-to-end work(S) Partial failure with no error, product value eroded
Steps per taskDetect abnormal trajectories(S) Cost and latency drift from lengthening
Loop detectionSpot repetition(S) A loop that burns the budget, seen on the bill
Tool failure rateIdentify failing toolsCostly retries, root cause not isolated
Latency per step and per taskDecompose the end-to-end pathDiagnosis impossible on long trajectories
Cost per task (cumulative tokens)Know the real cost of a task(S) Mispriced offering
Tool selection accuracyThe right tool at the right time(S) Degraded results with no signal
Human handoff rateMeasure real autonomyOperator load underestimated, ROI distorted
Guardrails, injection attemptsAgentic securityCommand injection undetected

Tools: Laminar and Langfuse (transcript, debug), Phoenix and Confident AI (multi-step eval), LangSmith and LangGraph Studio for a LangGraph stack.

flowchart TB
  P["Protocol layer: JSON-RPC over transport, sessions, schemas"] --> T["Execution layer: tool calls, latency, errors"] --> A["Task layer: impact on agent success"]
  A -.->|"cascade"| P
WhatWhy (decision)Impact if unmonitored
Latency and errors per tool and serverExecution health, isolationFailure mis-attributed to the agent
Session health, client-server desyncTransport stability(S) Desync that degrades the agent with no error
Authentication failures per identityDetect abuse and misconfigurationIllegitimate access, secret exposure
Enumeration / discovery spikesDetect hostile reconnaissance(S) Invisible reconnaissance before exploitation
Server mutationsDetect rug pull and schema poisoning(S) A trusted tool turned malicious
Tool pinning by hashTool integrity over time(S) Undetected modification
Invocation logRegulatory audit trailNo traceability, non-compliance
Availability per server, health checksGuarantee fallbackAgent looping on an unavailable server

Security: the OWASP MCP Top 10 reference (2025 edition, in beta) defines ten categories: MCP01 token mismanagement and secret exposure, MCP02 privilege escalation via scope creep, MCP03 tool poisoning, MCP04 software supply chain attacks and dependency tampering, MCP05 command injection and execution, MCP06 prompt injection via contextual payloads, MCP07 insufficient authentication and authorization, MCP08 lack of audit and telemetry, MCP09 shadow MCP servers, MCP10 context injection and over-sharing. Rug pull, tool shadowing and schema poisoning are attack techniques described in the literature, mostly falling under MCP03; they are not categories of the reference. The grid above directly addresses MCP08.

Tool call auto-approval removes the last barrier that does not depend on the model. The MCPTox benchmark (2025) ran 20 LLM agents against 1,348 tool poisoning cases built on 45 real MCP servers: the attack succeeds in 36.5% of cases on average, up to 72.8% for the most exposed model, and no model refuses more than 3% of attacks. The benchmark measures model susceptibility, not the effect of human approval; I still recommend human-in-the-loop validation on sensitive tools, because it does not rely on the judgment of the model under attack. Tools: mcp-scan (Invariant Labs / Snyk) and Cisco mcp-scanner (YARA), complementary, plus an MCP gateway with audit and OTel export.


Revised on 2 October 2026: official OWASP MCP Top 10 categories (2025 edition, in beta) and MCPTox benchmark figures in place of an unsourced claim about auto-approval.