Skip to main content

Fine-Tuning LLMs (Lite)

Learn when fine-tuning beats RAG and prompting, why LoRA is the default, and how to run, evaluate, and check a small LoRA fine-tune for forgetting.

~3 hours
12 Topics
Hands-on Scenarios

What You'll Learn

Understanding Why acme-assist Might Need Fine-Tuning

By now acme-assist answers from the help centre, with citations.

Deciding Whether You Need to Fine-Tune

Fine-tuning costs data, training time, and an artifact to maintain forever, so the default answer is no. This topic gives you the test to apply first.

Understanding What Fine-Tuning Changes

Fine-tuning moves a model's weights toward the examples you show it.

Understanding LoRA and QLoRA

LoRA is the technique you will actually use. A little intuition about it makes the three settings you tune easy to reason about.

Preparing a Small, Clean Dataset

The dataset is the real product of a fine-tune.

Training a LoRA Adapter

Training is a short, mostly mechanical step once the data is right. The goal here is to read the script until nothing in it surprises you.

Skills You'll Master

FINE-TUNINGLORAPEFTCATASTROPHIC-FORGETTINGAI-ENGINEERING

Curriculum Index12 topics

Career Impact

Roles that use the skills in this module.

  • AI Engineer

  • MLOps Engineer

  • Platform Engineer

See how this is asked in interviews

Practice on the Coding Sheet

Not a software engineer sheet. Every problem comes from real DevOps, SRE, Platform and Cloud interviews, from your first script to a system you build yourself.

Open the Coding Sheet

Frequently Asked Questions

Use RAG when the model is missing facts or the facts change often, because you can update a document without retraining. Fine-tune only when the facts are right but the tone, format, or style is wrong and prompting has already failed. Most problems that look like fine-tuning problems are really prompting or retrieval problems.

LoRA freezes the original model and trains a small add-on called an adapter. It needs far less memory than full fine-tuning, produces a file of megabytes instead of gigabytes, and leaves the base weights untouched. That makes it the sensible first choice for almost every team.

For a narrow style change, a few dozen clean, consistent examples are enough to see the effect, and a few hundred is a more realistic starting point for production. Quality matters more than volume, because every bad example teaches the wrong pattern. Always keep part of your data aside to test on.

Yes, for a very small model such as a 0.5 billion parameter one, but it is slow on a laptop CPU. A free Colab GPU finishes the same job in minutes, so use Colab as the default for the lab in this module.

It is the loss of general ability after training a model hard on a narrow dataset. LoRA lowers the risk because the base weights stay frozen, but it does not remove it. Test the tuned model on prompts unrelated to your task before you ship it.