How to choose a suitable AI model
The best model for you succeeds on your tasks at acceptable cost and latency.
What you will learn
- Explain: choose by task and capabilities
- Apply the idea in an example: How to choose a suitable AI model
- Recognize limitations and verify the exercise outcome
The best model for you succeeds on your tasks at acceptable cost and latency. A general leaderboard cannot replace local testing.
You can follow this course without an account or programming. Examples use synthetic data and AI outcomes need checking.
How it works, step by step
Choose by task and capabilities
Check language, context, tool calling, structured output and supported input modalities. Compare models on identical requests using the same rubric and a realistic budget. Temperature changes sampling, not truth. Reasoning effort allocates effort for compatible models and can increase time and cost; accepted values vary. Local models offer infrastructure control but require resources and maintenance. Open weights do not guarantee an unrestricted license or complete training code. Catalogs and prices change; read the current model card.
Define success before optimizing
Prepare representative requests and expected answers or properties. Include ordinary cases, ambiguity, missing data and tool errors. For RAG, measure evidence retrieval separately from answer correctness. Recall@k measures the share of relevant evidence found in the first k results; precision@k measures the share of returned results that are relevant. For agents, also track actions, permissions and stopping. An LLM-as-judge can speed evaluation but needs calibration against humans. Do not change prompt, model and index simultaneously if you want to understand what improved results.
Cost belongs to the whole path
Count model calls, input and output tokens, embeddings, reranking, vision and tools. Reasoning can consume billable tokens without an equally long visible response. Indexing costs and conversation costs differ. Use current prices rather than a permanent number from a course. A simple estimate is tokens/1,000,000 × price plus additional operations; application credits may have their own conversion. Compare cost per successful task, not just per call. Parallel execution can reduce latency while increasing total consumption. A budget needs a verified stopping condition.
The visual map
Follow the solid arrows for the main flow. Dashed blue arrows supply data or context; dashed pink arrows show feedback or returning results. Colors and shapes distinguish models, stores, decisions and outputs. On smaller screens, scroll horizontally to follow the entire diagram.
A complete example
Compare a small and larger model on ten Romanian questions, two tool calls and an illustrated document. Check vision and tool support separately.
Try it yourself
- Build a selection table and set thresholds before testing.
- Record the input, source and expected outcome before running the experiment. Use only the fictional data in the example.
- Follow the diagram stages. At every step record what information is received and produced; do not confuse intermediate output with the final outcome.
- Repeat after removing necessary information or making the input ambiguous. Check whether the system clarifies, stops or invents an answer.
- Compare with the explained solution. Keep the configuration, date, result and an explanation for differences. Change one thing and retest.
An explained solution
The comparison includes success, latency, cost and required capabilities; incompatible models are excluded before scoring. A successful exercise lets you show the connection between input, stages and outcome. When information is missing, a cautious answer is more useful than invented details. Compare more than style: check conditions, sources and operations too.
When it helps and what can go wrong
Open weights and free are not synonyms. Choose this approach when it improves a measured need. Keep a simple baseline and compare outcomes using identical inputs. One successful example does not establish reliability in every situation.
Check your understanding
Does zero temperature guarantee correctness?
No. It may reduce variation, but a stable answer can still be wrong.
Does a good style score prove agent success?
No. The agent may write beautifully after using the wrong tool or data.
What outcome should this exercise produce?
The comparison includes success, latency, cost and required capabilities; incompatible models are excluded before scoring.
Words to remember
- Reasoning effort: A model-dependent setting controlling reasoning effort.
- Regresie / Regression: A change that breaks a previously correct case.
- Latență / Latency: Time until a useful result.
Sources and your next step
To prepare: P03 — Temperature, limits and reasoning effort