What is an LLM and how does it answer?
An LLM builds answers token by token.
What you will learn
- Explain: from tokens to an answer
- Apply the idea in an example: What is an LLM and how does it answer?
- Recognize limitations and verify the exercise outcome
An LLM builds answers token by token. We can understand it without advanced formulas by following what it receives and generates.
You can follow this course without an account or programming. Examples use synthetic data and AI outcomes need checking.
How it works, step by step
From tokens to an answer
A token is a unit of text: sometimes a word, sometimes a word fragment or punctuation. The model processes context tokens and estimates a distribution over the next token. Selection repeats until stopping. Transformer attention combines information about relationships between tokens; it is not human attention. Training changes parameters, while inference uses the trained model. An answer can sound coherent without being true: generating text does not independently verify facts.
Context is the workbench
Context includes instructions, the current request, relevant history, retrieved documents and tool results. The context window has limited capacity; reserve room for the response too. More text does not automatically improve accuracy. Irrelevant information and contradictions can hide important evidence. Summarize history carefully and preserve the origin of facts. State describes the current run, a checkpoint saves a resumable point, and persistent memory retains information across runs. These are different mechanisms managed by the application.
Fluency and truth are checked separately
A hallucination is unsupported or incorrect content presented as an answer. It can arise from missing information, ambiguity or incorrectly combined patterns. Ask for evidence and check the original document, date, units and conditions. A real citation may still fail to support the claim. Use current sources for current facts, a calculator for arithmetic, and “cannot determine” when evidence is missing. Do not treat confidence expressed in prose as a calibrated probability.
The visual map
Follow the solid arrows for the main flow. Dashed blue arrows supply data or context; dashed pink arrows show feedback or returning results. Colors and shapes distinguish models, stores, decisions and outputs. On smaller screens, scroll horizontally to follow the entire diagram.
A complete example
“Summarize the policy in two sentences” becomes tokens. The model considers the instruction and policy, generates subsequent tokens and stops. It does not automatically search documents unless the application supplies retrieval.
Try it yourself
- Write a policy paragraph, request two summaries, then remove the policy.
- Record the input, source and expected outcome before running the experiment. Use only the fictional data in the example.
- Follow the diagram stages. At every step record what information is received and produced; do not confuse intermediate output with the final outcome.
- Repeat after removing necessary information or making the input ambiguous. Check whether the system clarifies, stops or invents an answer.
- Compare with the explained solution. Keep the configuration, date, result and an explanation for differences. Change one thing and retest.
An explained solution
The summary preserves the deadline and conditions. If no policy is supplied, the model should acknowledge that absence. A successful exercise lets you show the connection between input, stages and outcome. When information is missing, a cautious answer is more useful than invented details. Compare more than style: check conditions, sources and operations too.
When it helps and what can go wrong
“Most likely” does not mean “true”. Choose this approach when it improves a measured need. Keep a simple baseline and compare outcomes using identical inputs. One successful example does not establish reliability in every situation.
Check your understanding
Is a token always a word?
No; tokenization depends on the model and language. Use the appropriate tokenizer.
Is a checkpoint permanent user memory?
No. Saving a run does not imply retaining preferences across conversations.
What outcome should this exercise produce?
The summary preserves the deadline and conditions. If no policy is supplied, the model should acknowledge that absence.
Words to remember
- Token: A unit used by the model for text and context limits.
- Checkpoint: Saved state used to resume an execution.
- Grounding: Grounding an answer in verifiable evidence.
Sources and your next step
To prepare: F02 — How does a machine learn?