ML Infrastructure on AWS: SageMaker and Bedrock
Learn to choose, deploy, and operate ML infrastructure on AWS - from Rekognition to SageMaker training, deployment modes, drift detection, and Bedrock RAG.
What You'll Learn
Understanding Why ML Infrastructure Is a Separate Discipline
The model that runs perfectly and is silently wrong A payments team in Bengaluru gets a Slack message on a Friday evening.
Choosing Between Pre-Built AI Services and Custom Models
The pre-built services and what each one does The most expensive mistake in ML infrastructure is building a model AWS already built.
Understanding the SageMaker Training Workflow
What a training job actually is A training job is a short-lived, managed compute job.
Controlling Training Cost and Failures
Picking an instance without overspending Start with the smallest instance that finishes in a reasonable time.
Choosing the Right Deployment Mode for Inference
Why deployment is a separate decision from training Training produces a model artifact.
Detecting Model Drift with SageMaker Model Monitor
The two kinds of drift Return to the Friday-evening incident. The model never crashed. The world under it changed, and nobody was watching.
Skills You'll Master
Curriculum Index12 topics
Understanding Why ML Infrastructure Is a Separate Discipline
The model that runs perfectly and is silently wrong A payments team in Bengaluru gets a Slack message on a Friday...
Choosing Between Pre-Built AI Services and Custom Models
The pre-built services and what each one does The most expensive mistake in ML infrastructure is building a model AWS...
Understanding the SageMaker Training Workflow
What a training job actually is A training job is a short-lived, managed compute job.
Controlling Training Cost and Failures
Picking an instance without overspending Start with the smallest instance that finishes in a reasonable time.
Choosing the Right Deployment Mode for Inference
Why deployment is a separate decision from training Training produces a model artifact.
Detecting Model Drift with SageMaker Model Monitor
The two kinds of drift Return to the Friday-evening incident. The model never crashed.
Managing Features with SageMaker Feature Store
Training-serving skew in plain terms Model Monitor catches drift once a model is live.
Automating Retraining with SageMaker Pipelines
What a pipeline chains together Models do not all decay at the same speed.
Using Amazon Bedrock When You Should Not Train a Model
Foundation models versus custom training Many new ML projects no longer start with training.
Controlling Access and Cost for ML Workloads
Giving SageMaker a least-privilege execution role Every training job, endpoint, and monitoring job assumes an execution...
Hands-On Lab: Train, Deploy, and Monitor an XGBoost Model
Before you start You will build a small fraud model on synthetic data, deploy it with data capture turned on, set up...
Quick Reference and Common Mistakes
Quick reference Common mistakes Teams reach for SageMaker before checking whether Rekognition, Textract, Comprehend, or...
Career Impact
Roles that use the skills in this module.
AI Engineer
MLOps Engineer
Platform Engineer
Cloud Engineer
Next Modules
Practice on the Coding Sheet
Not a software engineer sheet. Every problem comes from real DevOps, SRE, Platform and Cloud interviews, from your first script to a system you build yourself.
Open the Coding SheetFrequently Asked Questions
Use a pre-built service such as Rekognition, Textract, Comprehend, or Transcribe when the task is generic and thousands of other companies need the same thing. Move to SageMaker only when the prediction depends on your own proprietary data, such as your fraud patterns or your customers. A single API call that ships this week usually beats a training project that takes months.
Training jobs bill only while they run, and batch transform bills only for the job. Real-time endpoints are different: they bill every hour they exist, even with zero traffic. Notebook instances and Studio apps can also keep billing if left running, so always check for leftovers.
A real-time endpoint keeps instances running at all times, so responses are fast and you pay for uptime. Serverless inference provisions compute when a request arrives and scales to zero when idle, so you pay per use but the first request after a quiet period is slower. Pick based on how steady your traffic is.
Any model that serves live decisions for more than a few weeks should be monitored, because drift is silent. A model that is only used for a one-off batch analysis does not need it. Start with data quality monitoring and add model quality monitoring once you can collect ground-truth labels.
No. Bedrock gives you managed access to foundation models for language and reasoning tasks, while SageMaker is for training and serving models on your own structured data. Many production systems use both: SageMaker for a fraud score, Bedrock for summarising the case notes.