Temperature, limits and reasoning effort

Model settings have tradeoffs.

What you will learn

  • Explain: choose by task and capabilities
  • Apply the idea in an example: Temperature, limits and reasoning effort
  • Recognize limitations and verify the exercise outcome

Model settings have tradeoffs. Understand each one before changing all values together.

You can follow this course without an account or programming. Examples use synthetic data and AI outcomes need checking.

How it works, step by step

Choose by task and capabilities

Check language, context, tool calling, structured output and supported input modalities. Compare models on identical requests using the same rubric and a realistic budget. Temperature changes sampling, not truth. Reasoning effort allocates effort for compatible models and can increase time and cost; accepted values vary. Local models offer infrastructure control but require resources and maintenance. Open weights do not guarantee an unrestricted license or complete training code. Catalogs and prices change; read the current model card.

From tokens to an answer

A token is a unit of text: sometimes a word, sometimes a word fragment or punctuation. The model processes context tokens and estimates a distribution over the next token. Selection repeats until stopping. Transformer attention combines information about relationships between tokens; it is not human attention. Training changes parameters, while inference uses the trained model. An answer can sound coherent without being true: generating text does not independently verify facts.

Cost belongs to the whole path

Count model calls, input and output tokens, embeddings, reranking, vision and tools. Reasoning can consume billable tokens without an equally long visible response. Indexing costs and conversation costs differ. Use current prices rather than a permanent number from a course. A simple estimate is tokens/1,000,000 × price plus additional operations; application credits may have their own conversion. Compare cost per successful task, not just per call. Parallel execution can reduce latency while increasing total consumption. A budget needs a verified stopping condition.

The visual map

Temperature, limits and reasoning effort Follow the solid arrows for the main flow. Dashed blue arrows supply data or context; dashed pink arrows show feedback or returning results. Colors and shapes distinguish models, stores, decisions and outputs. Connections: User request → Agent / LLM; Agent / LLM → Final answer; Final answer → Quality measurements; Sampling settings → Agent / LLM; Reasoning effort → Agent / LLM; Token and cost budget → Agent / LLM; Calls and token consumption → Quality measurements. P03 · Relationship map Temperature, limits and reasoning effort Input User request Decision / control Sampling settings Decision / control Reasoning effort AI model / agent Agent / LLM Decision / control Token and cost budget Output Final answer Data Calls and token consumption Decision / control Quality measurements Main flow Data and context Feedback and return

Follow the solid arrows for the main flow. Dashed blue arrows supply data or context; dashed pink arrows show feedback or returning results. Colors and shapes distinguish models, stores, decisions and outputs. On smaller screens, scroll horizontally to follow the entire diagram.

A complete example

Classification has strict criteria; brainstorming needs variety. Test temperatures where supported and compare results. Compare effort separately on complex tasks.

Try it yourself

  1. Run five requests with two settings and record quality, time and consumption.
  2. Record the input, source and expected outcome before running the experiment. Use only the fictional data in the example.
  3. Follow the diagram stages. At every step record what information is received and produced; do not confuse intermediate output with the final outcome.
  4. Repeat after removing necessary information or making the input ambiguous. Check whether the system clarifies, stops or invents an answer.
  5. Compare with the explained solution. Keep the configuration, date, result and an explanation for differences. Change one thing and retest.

An explained solution

Choose a value meeting the criteria at an acceptable cost; keep the model and other settings fixed. A successful exercise lets you show the connection between input, stages and outcome. When information is missing, a cautious answer is more useful than invented details. Compare more than style: check conditions, sources and operations too.

When it helps and what can go wrong

Some models do not support identical parameters or values. Choose this approach when it improves a measured need. Keep a simple baseline and compare outcomes using identical inputs. One successful example does not establish reliability in every situation.

Check your understanding

Does zero temperature guarantee correctness?

No. It may reduce variation, but a stable answer can still be wrong.

Is a token always a word?

No; tokenization depends on the model and language. Use the appropriate tokenizer.

What outcome should this exercise produce?

Choose a value meeting the criteria at an acceptable cost; keep the model and other settings fixed.

Words to remember

  • Reasoning effort: A model-dependent setting controlling reasoning effort.
  • Token: A unit used by the model for text and context limits.
  • Latență / Latency: Time until a useful result.

Sources and your next step

To prepare: P02 — Context engineering: what reaches the model?