Playground as a laboratory

Playground makes conversation steps visible.

What you will learn

  • Explain: a bounded agent loop
  • Apply the idea in an example: Playground as a laboratory
  • Recognize limitations and verify the exercise outcome

Playground makes conversation steps visible. Use it for controlled comparisons rather than only successful demos.

These steps refer to features identified in the application code. This material has not been validated on a live instance; check the documentation and the options in your version.

How it works, step by step

A bounded agent loop

An agent combines a model with tools, a goal and state. It observes the request, chooses a step, receives a result and decides whether to continue. A chatbot may only generate text; an agent may request an external action. Commercial definitions vary, so inspect actual control. Define success and stopping: a verified result, a maximum step count, repeated errors or an exhausted budget. Use narrow capabilities and minimal permissions. Useful autonomy means freedom within boundaries enforced by software.

Define success before optimizing

Prepare representative requests and expected answers or properties. Include ordinary cases, ambiguity, missing data and tool errors. For RAG, measure evidence retrieval separately from answer correctness. Recall@k measures the share of relevant evidence found in the first k results; precision@k measures the share of returned results that are relevant. For agents, also track actions, permissions and stopping. An LLM-as-judge can speed evaluation but needs calibration against humans. Do not change prompt, model and index simultaneously if you want to understand what improved results.

Errors need explicit paths

A timeout differs from a permission error. Retry only transient errors, with bounded attempts and increasing delays. An operation with side effects may execute before its response is lost; retrying without an idempotency key can duplicate the effect. Streaming shows progress but does not guarantee completion. Checkpoints may allow failed stages to resume within runtime limits. Store the run identifier, stage, duration and error while removing secrets from logs. Measure slow paths as well as averages.

The visual map

Playground as a laboratory Follow the solid arrows for the main flow. Dashed blue arrows supply data or context; dashed pink arrows show feedback or returning results. Colors and shapes distinguish models, stores, decisions and outputs. Connections: Representative test questions → Playground tests; Playground tests → Streaming events; Streaming events → Result to verify; Result to verify → Compare against expected results; Configured agents → Playground tests; Multi-agent workflow → Playground tests; Execution trace → Compare against expected results; User feedback → Compare against expected results. U02 · Relationship map Playground as a laboratory Input Representative test questions AI model / agent Configured agents Processing Multi-agent workflow Processing Playground tests Data Streaming events Store / index Execution trace Output Result to verify Data User feedback Decision / control Compare against expected results Main flow Data and context Feedback and return

Follow the solid arrows for the main flow. Dashed blue arrows supply data or context; dashed pink arrows show feedback or returning results. Colors and shapes distinguish models, stores, decisions and outputs. On smaller screens, scroll horizontally to follow the entire diagram.

A complete example

Choose an agent or workflow, submit a question and inspect streaming and steps. For files, check upload. Use fresh conversations in comparisons so history does not change the test.

Try it yourself

  1. Run the same question with two configurations and record differences.
  2. Record the input, source and expected outcome before running the experiment. Use only the fictional data in the example.
  3. Follow the diagram stages. At every step record what information is received and produced; do not confuse intermediate output with the final outcome.
  4. Repeat after removing necessary information or making the input ambiguous. Check whether the system clarifies, stops or invents an answer.
  5. Compare with the explained solution. Keep the configuration, date, result and an explanation for differences. Change one thing and retest.

An explained solution

The report includes request, target, steps and outcome, not just final text. A successful exercise lets you show the connection between input, stages and outcome. When information is missing, a cautious answer is more useful than invented details. Compare more than style: check conditions, sources and operations too.

When it helps and what can go wrong

A conversation attachment does not automatically become permanent RAG data. Choose this approach when it improves a measured need. Keep a simple baseline and compare outcomes using identical inputs. One successful example does not establish reliability in every situation.

Check your understanding

Does more autonomy always give a better result?

No. It also increases room for errors, costs and unnecessary actions.

Does a good style score prove agent success?

No. The agent may write beautifully after using the wrong tool or data.

What outcome should this exercise produce?

The report includes request, target, steps and outcome, not just final text.

Words to remember

  • Agent: A system that can choose steps and use tools toward a goal.
  • Regresie / Regression: A change that breaks a previously correct case.
  • Idempotency: Controlled repetition without duplicating an effect.

Sources and your next step

To prepare: U01 — A map of the application