Optimization and operations
Optimize an observed need: latency, cost or failures.
What you will learn
- Explain: errors need explicit paths
- Apply the idea in an example: Optimization and operations
- Recognize limitations and verify the exercise outcome
Optimize an observed need: latency, cost or failures. More complicated configurations are not automatically more robust.
This is a conceptual or external lab. It does not assume the application exposes every control described.
How it works, step by step
Errors need explicit paths
A timeout differs from a permission error. Retry only transient errors, with bounded attempts and increasing delays. An operation with side effects may execute before its response is lost; retrying without an idempotency key can duplicate the effect. Streaming shows progress but does not guarantee completion. Checkpoints may allow failed stages to resume within runtime limits. Store the run identifier, stage, duration and error while removing secrets from logs. Measure slow paths as well as averages.
Cost belongs to the whole path
Count model calls, input and output tokens, embeddings, reranking, vision and tools. Reasoning can consume billable tokens without an equally long visible response. Indexing costs and conversation costs differ. Use current prices rather than a permanent number from a course. A simple estimate is tokens/1,000,000 × price plus additional operations; application credits may have their own conversion. Compare cost per successful task, not just per call. Parallel execution can reduce latency while increasing total consumption. A budget needs a verified stopping condition.
Define success before optimizing
Prepare representative requests and expected answers or properties. Include ordinary cases, ambiguity, missing data and tool errors. For RAG, measure evidence retrieval separately from answer correctness. Recall@k measures the share of relevant evidence found in the first k results; precision@k measures the share of returned results that are relevant. For agents, also track actions, permissions and stopping. An LLM-as-judge can speed evaluation but needs calibration against humans. Do not change prompt, model and index simultaneously if you want to understand what improved results.
The visual map
Follow the solid arrows for the main flow. Dashed blue arrows supply data or context; dashed pink arrows show feedback or returning results. Colors and shapes distinguish models, stores, decisions and outputs. On smaller screens, scroll horizontally to follow the entire diagram.
A complete example
A policy cache may reduce latency, but live stock needs rapid expiry. Retry helps transient timeout, not an invalid key. Fallback must preserve conditions and access.
Try it yourself
- Design bounded retry, cache expiry and conceptual fallback.
- Record the input, source and expected outcome before running the experiment. Use only the fictional data in the example.
- Follow the diagram stages. At every step record what information is received and produced; do not confuse intermediate output with the final outcome.
- Repeat after removing necessary information or making the input ambiguous. Check whether the system clarifies, stops or invents an answer.
- Compare with the explained solution. Keep the configuration, date, result and an explanation for differences. Change one thing and retest.
An explained solution
Measure before and after, include failures and check for regressions. A successful exercise lets you show the connection between input, stages and outcome. When information is missing, a cautious answer is more useful than invented details. Compare more than style: check conditions, sources and operations too.
When it helps and what can go wrong
Do not assume the product exposes all these options in its interface. Choose this approach when it improves a measured need. Keep a simple baseline and compare outcomes using identical inputs. One successful example does not establish reliability in every situation.
Check your understanding
Does a timeout prove the operation did not execute?
No. Check state before retrying an action with side effects.
Are more simultaneous calls automatically cheaper?
No. Elapsed time and total consumption are different quantities.
What outcome should this exercise produce?
Measure before and after, include failures and check for regressions.
Words to remember
- Idempotency: Controlled repetition without duplicating an effect.
- Latență / Latency: Time until a useful result.
- Regresie / Regression: A change that breaks a previously correct case.
Sources and your next step
To prepare: E05 — Tokens, credits, budgets and total cost