Evaluating agents and teams

An agent can write a good answer after a wrong action.

What you will learn

  • Explain: define success before optimizing
  • Apply the idea in an example: Evaluating agents and teams
  • Recognize limitations and verify the exercise outcome

An agent can write a good answer after a wrong action. Evaluate the whole path.

You can follow this course without an account or programming. Examples use synthetic data and AI outcomes need checking.

How it works, step by step

Define success before optimizing

Prepare representative requests and expected answers or properties. Include ordinary cases, ambiguity, missing data and tool errors. For RAG, measure evidence retrieval separately from answer correctness. Recall@k measures the share of relevant evidence found in the first k results; precision@k measures the share of returned results that are relevant. For agents, also track actions, permissions and stopping. An LLM-as-judge can speed evaluation but needs calibration against humans. Do not change prompt, model and index simultaneously if you want to understand what improved results.

The model proposes, the tool executes

Tool calling produces a call request with a name and arguments. The application validates the schema, permissions and limits before execution. The result returns to context for the next step. For example, get_product_stock takes a SKU and returns a quantity; its description should explain that it does not reserve products. Distinguish empty results, validation errors, missing access and timeouts. Valid JSON does not prove the arguments are correct or the action is authorized. For external effects, also consider duplicate execution.

A workflow has explicit structure

A node is a stage, an edge is a transition, and state carries results between stages. A workflow has an application-defined path, although the model may choose some branches. Agentic systems include both workflows and dynamically acting agents. Nymrio offers ReAct, Reflection, Sequential, Orchestrator, Router, Parallel and Plan-and-Execute. Each changes coordination; it does not automatically make a model an expert. More agents mean more calls and contracts between roles. Start with the simplest solution that passes your tests.

The visual map

Evaluating agents and teams Follow the solid arrows for the main flow. Dashed blue arrows supply data or context; dashed pink arrows show feedback or returning results. Colors and shapes distinguish models, stores, decisions and outputs. Connections: Evaluation dataset → Agent / LLM; Agent / LLM → Task success and safe actions; Task success and safe actions → Calibrated judge + human review; Calibrated judge + human review → Human verification; Evaluation dataset → Multi-agent workflow; Multi-agent workflow → Task success and safe actions; Expected facts and sources → Task success and safe actions; Access permissions → Task success and safe actions; Execution trace → Calibrated judge + human review. E03 · Relationship map Evaluating agents and teams Store / index Evaluation dataset Data Expected facts and sources Decision / control Access permissions AI model / agent Agent / LLM Processing Multi-agent workflow Store / index Execution trace Decision / control Task success and safe actions Decision / control Calibrated judge + human review Decision / control Human verification Main flow Data and context Feedback and return

Follow the solid arrows for the main flow. Dashed blue arrows supply data or context; dashed pink arrows show feedback or returning results. Colors and shapes distinguish models, stores, decisions and outputs. On smaller screens, scroll horizontally to follow the entire diagram.

A complete example

An agent reports the right quantity but queried another SKU. A team cites policy but omits worker disagreement. Both are failures despite convincing final text.

Try it yourself

  1. Compare two traces and mark the first deviation from the goal.
  2. Record the input, source and expected outcome before running the experiment. Use only the fictional data in the example.
  3. Follow the diagram stages. At every step record what information is received and produced; do not confuse intermediate output with the final outcome.
  4. Repeat after removing necessary information or making the input ambiguous. Check whether the system clarifies, stops or invents an answer.
  5. Compare with the explained solution. Keep the configuration, date, result and an explanation for differences. Change one thing and retest.

An explained solution

Check arguments, roles, operations, sources and stopping, then text quality. A successful exercise lets you show the connection between input, stages and outcome. When information is missing, a cautious answer is more useful than invented details. Compare more than style: check conditions, sources and operations too.

When it helps and what can go wrong

LLM-as-judge can be biased and needs human calibration. Choose this approach when it improves a measured need. Keep a simple baseline and compare outcomes using identical inputs. One successful example does not establish reliability in every situation.

Check your understanding

Does a good style score prove agent success?

No. The agent may write beautifully after using the wrong tool or data.

Who executes a call proposed by the model?

The application or host service after validation. The model does not automatically gain unrestricted access.

What outcome should this exercise produce?

Check arguments, roles, operations, sources and stopping, then text quality.

Words to remember

  • Regresie / Regression: A change that breaks a previously correct case.
  • Tool calling: A structured request to use an external capability.
  • Nod: A stage in an execution graph.

Sources and your next step

To prepare: E02 — Evaluating RAG stage by stage