Observability and debugging

Observability connects outcomes with the steps producing them.

What you will learn

  • Explain: errors need explicit paths
  • Apply the idea in an example: Observability and debugging
  • Recognize limitations and verify the exercise outcome

Observability connects outcomes with the steps producing them. Diagnostic data should help without exposing secrets.

You can follow this course without an account or programming. Examples use synthetic data and AI outcomes need checking.

How it works, step by step

Errors need explicit paths

A timeout differs from a permission error. Retry only transient errors, with bounded attempts and increasing delays. An operation with side effects may execute before its response is lost; retrying without an idempotency key can duplicate the effect. Streaming shows progress but does not guarantee completion. Checkpoints may allow failed stages to resume within runtime limits. Store the run identifier, stage, duration and error while removing secrets from logs. Measure slow paths as well as averages.

Define success before optimizing

Prepare representative requests and expected answers or properties. Include ordinary cases, ambiguity, missing data and tool errors. For RAG, measure evidence retrieval separately from answer correctness. Recall@k measures the share of relevant evidence found in the first k results; precision@k measures the share of returned results that are relevant. For agents, also track actions, permissions and stopping. An LLM-as-judge can speed evaluation but needs calibration against humans. Do not change prompt, model and index simultaneously if you want to understand what improved results.

Check access before use

Permissions must be enforced by servers and the data layer. Do not let the model decide whether a user may see a document. Filter sources before they enter context, not after generating the answer. Use minimal-access credentials, keep them out of prompts, screenshots and browser code, and rotate exposed keys. A widget domain allowlist does not replace API authentication. Authentication identifies you; authorization determines the operations and data you can use. Test denied access too.

The visual map

Observability and debugging Follow the solid arrows for the main flow. Dashed blue arrows supply data or context; dashed pink arrows show feedback or returning results. Colors and shapes distinguish models, stores, decisions and outputs. Connections: User request → Agent / LLM; Agent / LLM → Execution trace; Execution trace → Steps, errors and timing; Steps, errors and timing → Compare against expected results; Compare against expected results → User revises the request; User revises the request → Playground tests; Multi-agent workflow → Execution trace; User revises the request → Agent / LLM. E04 · Relationship map Observability and debugging Input User request AI model / agent Agent / LLM Processing Multi-agent workflow Store / index Execution trace Data Steps, errors and timing Decision / control Compare against expected results Input User revises the request Processing Playground tests Main flow Data and context Feedback and return

Follow the solid arrows for the main flow. Dashed blue arrows supply data or context; dashed pink arrows show feedback or returning results. Colors and shapes distinguish models, stores, decisions and outputs. On smaller screens, scroll horizontally to follow the entire diagram.

A complete example

A slow response has retrieval 0.2 s, rerank 0.4 s and external tool 8 s. A run identifier connects stages; prompt and model versions are recorded.

Try it yourself

  1. Build a timeline with stage, duration, outcome and run identifier.
  2. Record the input, source and expected outcome before running the experiment. Use only the fictional data in the example.
  3. Follow the diagram stages. At every step record what information is received and produced; do not confuse intermediate output with the final outcome.
  4. Repeat after removing necessary information or making the input ambiguous. Check whether the system clarifies, stops or invents an answer.
  5. Compare with the explained solution. Keep the configuration, date, result and an explanation for differences. Change one thing and retest.

An explained solution

Identify the tool as the latency source and retain error context without credentials. A successful exercise lets you show the connection between input, stages and outcome. When information is missing, a cautious answer is more useful than invented details. Compare more than style: check conditions, sources and operations too.

When it helps and what can go wrong

External tracing tools are not assumed UI controls. Choose this approach when it improves a measured need. Keep a simple baseline and compare outcomes using identical inputs. One successful example does not establish reliability in every situation.

Check your understanding

Does a timeout prove the operation did not execute?

No. Check state before retrying an action with side effects.

Does a good style score prove agent success?

No. The agent may write beautifully after using the wrong tool or data.

What outcome should this exercise produce?

Identify the tool as the latency source and retain error context without credentials.

Words to remember

  • Idempotency: Controlled repetition without duplicating an effect.
  • Regresie / Regression: A change that breaks a previously correct case.
  • Autorizare / Authorization: Checking the right to access an operation or resource.

Sources and your next step

To prepare: E03 — Evaluating agents and teams