Skip to main content

LLM Fundamentals, Model Selection, and Prompt Engineering

Learn how LLMs generate text, how to choose a model on capability, cost, and latency, and how to call and prompt it for reliable structured output.

~4 hours
12 Topics
Hands-on Scenarios

What You'll Learn

Understanding Why This Module Comes Before Everything Else

You have a local model running and a helper to call it.

Understanding What a Large Language Model Actually Is

A large language model (LLM) is a next-token predictor trained at huge scale.

Choosing the Right Model for the Job

Picking a model is an engineering tradeoff, not a leaderboard lookup. The strongest model is rarely the right default for a production feature.

Understanding Inference: How a Response Actually Gets Generated

A model does not compose a whole answer at once.

Working With the Request and Response Pattern

Almost every LLM call you make is a list of role-tagged messages in and a message out.

Controlling Output With Temperature and Sampling

Temperature decides how adventurous the model is when picking the next token, which makes it the cheapest lever for consistency.

Skills You'll Master

LLMMODEL-SELECTIONPROMPT-ENGINEERINGCONTEXT-WINDOWSSTRUCTURED-OUTPUT

Curriculum Index12 topics

Career Impact

Roles that use the skills in this module.

  • AI Engineer

  • MLOps Engineer

  • Platform Engineer

See how this is asked in interviews

Practice on the Coding Sheet

Not a software engineer sheet. Every problem comes from real DevOps, SRE, Platform and Cloud interviews, from your first script to a system you build yourself.

Open the Coding Sheet

Frequently Asked Questions

A token is a sub-word chunk of text from the model's fixed vocabulary, and it is the unit models read, generate, and bill by. As a rough starting estimate, one token is around three-quarters of an English word, but the ratio varies by tokenizer and language. Measure real counts with your provider's tokenizer or the usage fields in the API response.

No. The strongest model costs more and responds more slowly, and many tasks such as classification and extraction do not need it. Pick the cheapest model that reliably clears the quality bar on an evaluation set built from your own data.

Models sample from a probability distribution, and higher temperature spreads that sampling more widely. Even at temperature 0, output is not guaranteed to be identical across calls because of server-side batching and floating point effects. Design systems that tolerate small variation.

No. A system prompt sets behavior but a determined user can sometimes get a model to ignore it. Enforce authorization, tool permissions, and data access in your application code and infrastructure.