How to measure whether AI is useful

Usefulness is measured in user outcomes.

What you will learn

  • Explain: define success before optimizing
  • Apply the idea in an example: How to measure whether AI is useful
  • Recognize limitations and verify the exercise outcome

Usefulness is measured in user outcomes. Define criteria before seeing answers rather than changing rules afterward.

You can follow this course without an account or programming. Examples use synthetic data and AI outcomes need checking.

How it works, step by step

Define success before optimizing

Prepare representative requests and expected answers or properties. Include ordinary cases, ambiguity, missing data and tool errors. For RAG, measure evidence retrieval separately from answer correctness. Recall@k measures the share of relevant evidence found in the first k results; precision@k measures the share of returned results that are relevant. For agents, also track actions, permissions and stopping. An LLM-as-judge can speed evaluation but needs calibration against humans. Do not change prompt, model and index simultaneously if you want to understand what improved results.

Fluency and truth are checked separately

A hallucination is unsupported or incorrect content presented as an answer. It can arise from missing information, ambiguity or incorrectly combined patterns. Ask for evidence and check the original document, date, units and conditions. A real citation may still fail to support the claim. Use current sources for current facts, a calculator for arithmetic, and “cannot determine” when evidence is missing. Do not treat confidence expressed in prose as a calibrated probability.

Cost belongs to the whole path

Count model calls, input and output tokens, embeddings, reranking, vision and tools. Reasoning can consume billable tokens without an equally long visible response. Indexing costs and conversation costs differ. Use current prices rather than a permanent number from a course. A simple estimate is tokens/1,000,000 × price plus additional operations; application credits may have their own conversion. Compare cost per successful task, not just per call. Parallel execution can reduce latency while increasing total consumption. A budget needs a verified stopping condition.

The visual map

How to measure whether AI is useful Follow the solid arrows for the main flow. Dashed blue arrows supply data or context; dashed pink arrows show feedback or returning results. Colors and shapes distinguish models, stores, decisions and outputs. Connections: Problem and success criteria → Agent / LLM; Agent / LLM → Compare against expected results; Compare against expected results → Quality measurements; Quality measurements → Human verification; Evaluation dataset → Agent / LLM; Expected facts and sources → Compare against expected results; Simple baseline → Compare against expected results. E01 · Relationship map How to measure whether AI is useful Input Problem and success criteria Store / index Evaluation dataset Data Expected facts and sources Processing Simple baseline AI model / agent Agent / LLM Decision / control Compare against expected results Decision / control Quality measurements Decision / control Human verification Main flow Data and context Feedback and return

Follow the solid arrows for the main flow. Dashed blue arrows supply data or context; dashed pink arrows show feedback or returning results. Colors and shapes distinguish models, stores, decisions and outputs. On smaller screens, scroll horizontally to follow the entire diagram.

A complete example

For FAQ score correct deadline, preserved conditions, clarity and absence of invented claims. Elegant prose with a wrong deadline fails the critical criterion.

Try it yourself

  1. Build a rubric for 20 requests and specify which failures block launch.
  2. Record the input, source and expected outcome before running the experiment. Use only the fictional data in the example.
  3. Follow the diagram stages. At every step record what information is received and produced; do not confuse intermediate output with the final outcome.
  4. Repeat after removing necessary information or making the input ambiguous. Check whether the system clarifies, stops or invents an answer.
  5. Compare with the explained solution. Keep the configuration, date, result and an explanation for differences. Change one thing and retest.

An explained solution

The rubric separates required criteria from style preferences and allows repeatable comparisons. A successful exercise lets you show the connection between input, stages and outcome. When information is missing, a cautious answer is more useful than invented details. Compare more than style: check conditions, sources and operations too.

When it helps and what can go wrong

An average can hide a rare but important failure. Choose this approach when it improves a measured need. Keep a simple baseline and compare outcomes using identical inputs. One successful example does not establish reliability in every situation.

Check your understanding

Does a good style score prove agent success?

No. The agent may write beautifully after using the wrong tool or data.

Is a link in an answer enough?

No; it must support the specific claim and fit the time and circumstances.

What outcome should this exercise produce?

The rubric separates required criteria from style preferences and allows repeatable comparisons.

Words to remember

  • Regresie / Regression: A change that breaks a previously correct case.
  • Grounding: Grounding an answer in verifiable evidence.
  • Latență / Latency: Time until a useful result.

Sources and your next step

To prepare: U02 — Playground as a laboratory · D04 — Good questions for testing RAG