Machine Learning for AI Engineers
Learn classical ML for AI engineers: regression, classification, tree models, honest validation, and metrics that tell you when not to use an LLM.
What You'll Learn
Understanding Why AI Engineers Still Need Classical ML
acme-assist now answers questions from the help centre. Next door, acme-shop's risk team has a different problem: scoring card payments for fraud.
Choosing the Right Kind of Learning
Before you pick an algorithm, decide what question you are asking. Choosing the wrong kind of learning wastes days.
Building Your First Regression Model
Regression models learn a relationship between input features and a number. Start with the simplest one, then compare stronger ones against it.
Preprocessing Data With a Pipeline
Raw data almost never goes straight into a model.
Validating a Model and Diagnosing Overfitting
A single train and test split can get lucky or unlucky.
Measuring Regression Performance
Regression metrics measure how far predictions land from the real values. Choose the metric that matches what a wrong prediction costs.
Skills You'll Master
Curriculum Index14 topics
Understanding Why AI Engineers Still Need Classical ML
acme-assist now answers questions from the help centre.
Choosing the Right Kind of Learning
Before you pick an algorithm, decide what question you are asking. Choosing the wrong kind of learning wastes days.
Building Your First Regression Model
Regression models learn a relationship between input features and a number.
Preprocessing Data With a Pipeline
Raw data almost never goes straight into a model.
Validating a Model and Diagnosing Overfitting
A single train and test split can get lucky or unlucky.
Measuring Regression Performance
Regression metrics measure how far predictions land from the real values.
Measuring Classification Beyond Accuracy
This is the topic that would have saved the fraud model from the start of the module.
Choosing a Threshold and Handling Imbalance
Most classifiers output a probability, and a threshold turns it into a label.
Recognizing Data Leakage
You met one form of leakage already: fitting preprocessing on the full dataset.
Tuning Hyperparameters and Knowing What to Skip
Tuning improves a model without touching the test set.
Hands-On Lab: Fraud Classification With Proper Evaluation
You will train two models on the kit's transactions data, judge them with the right metrics, then break the evaluation...
Quick Reference
A compact map of which tool or metric to reach for. Decision table Habits to repeat on every project
Common Mistakes
Seven mistakes cause most wasted weeks, each with its fix.
Reviewing What You Built and What Comes Next
You can now pick between regression and classification, build a leak-free pipeline, judge a model on the right metric...
Career Impact
Roles that use the skills in this module.
AI Engineer
MLOps Engineer
Platform Engineer
Next Modules
Related Guides
Practice on the Coding Sheet
Not a software engineer sheet. Every problem comes from real DevOps, SRE, Platform and Cloud interviews, from your first script to a system you build yourself.
Open the Coding SheetFrequently Asked Questions
Use classical ML when the data is structured and tabular, the task repeats, and you have labeled history, such as fraud scoring, churn, or demand prediction. It is usually faster, cheaper, and easier to evaluate than an LLM for those jobs. Use an LLM when the input is free text or the task needs language understanding.
Fraud is rare, so a model that always predicts not fraud can score above 99 percent accuracy while catching nothing. Precision, recall, and PR-AUC on the fraud class show whether the model finds the cases you care about.
Leakage is when information that would not exist at prediction time sneaks into training or evaluation. The model then looks excellent in testing and fails in production. Always ask whether a feature's value would be known at the moment a real prediction is made.
No. Tree-based models compare values inside one feature at a time, so scaling does not change their splits. Scaling matters for models such as logistic regression, K-Means, and SVMs.