Skip to main content

ML Infrastructure on AWS: SageMaker and Bedrock

Learn to choose, deploy, and operate ML infrastructure on AWS - from Rekognition to SageMaker training, deployment modes, drift detection, and Bedrock RAG.

~3.5 hours
12 Topics
Hands-on Scenarios

What You'll Learn

Understanding Why ML Infrastructure Is a Separate Discipline

The model that runs perfectly and is silently wrong A payments team in Bengaluru gets a Slack message on a Friday evening.

Choosing Between Pre-Built AI Services and Custom Models

The pre-built services and what each one does The most expensive mistake in ML infrastructure is building a model AWS already built.

Understanding the SageMaker Training Workflow

What a training job actually is A training job is a short-lived, managed compute job.

Controlling Training Cost and Failures

Picking an instance without overspending Start with the smallest instance that finishes in a reasonable time.

Choosing the Right Deployment Mode for Inference

Why deployment is a separate decision from training Training produces a model artifact.

Detecting Model Drift with SageMaker Model Monitor

The two kinds of drift Return to the Friday-evening incident. The model never crashed. The world under it changed, and nobody was watching.

Skills You'll Master

SAGEMAKERBEDROCKMACHINE-LEARNINGMODEL-DEPLOYMENTMLOPS

Curriculum Index12 topics

1

Understanding Why ML Infrastructure Is a Separate Discipline

The model that runs perfectly and is silently wrong A payments team in Bengaluru gets a Slack message on a Friday...

2

Choosing Between Pre-Built AI Services and Custom Models

The pre-built services and what each one does The most expensive mistake in ML infrastructure is building a model AWS...

3

Understanding the SageMaker Training Workflow

What a training job actually is A training job is a short-lived, managed compute job.

4

Controlling Training Cost and Failures

Picking an instance without overspending Start with the smallest instance that finishes in a reasonable time.

5

Choosing the Right Deployment Mode for Inference

Why deployment is a separate decision from training Training produces a model artifact.

6

Detecting Model Drift with SageMaker Model Monitor

The two kinds of drift Return to the Friday-evening incident. The model never crashed.

7

Managing Features with SageMaker Feature Store

Training-serving skew in plain terms Model Monitor catches drift once a model is live.

8

Automating Retraining with SageMaker Pipelines

What a pipeline chains together Models do not all decay at the same speed.

9

Using Amazon Bedrock When You Should Not Train a Model

Foundation models versus custom training Many new ML projects no longer start with training.

10

Controlling Access and Cost for ML Workloads

Giving SageMaker a least-privilege execution role Every training job, endpoint, and monitoring job assumes an execution...

11

Hands-On Lab: Train, Deploy, and Monitor an XGBoost Model

Before you start You will build a small fraud model on synthetic data, deploy it with data capture turned on, set up...

12

Quick Reference and Common Mistakes

Quick reference Common mistakes Teams reach for SageMaker before checking whether Rekognition, Textract, Comprehend, or...

Career Impact

Roles that use the skills in this module.

  • AI Engineer

  • MLOps Engineer

  • Platform Engineer

  • Cloud Engineer

See how this is asked in interviews

Practice on the Coding Sheet

Not a software engineer sheet. Every problem comes from real DevOps, SRE, Platform and Cloud interviews, from your first script to a system you build yourself.

Open the Coding Sheet

Frequently Asked Questions

Use a pre-built service such as Rekognition, Textract, Comprehend, or Transcribe when the task is generic and thousands of other companies need the same thing. Move to SageMaker only when the prediction depends on your own proprietary data, such as your fraud patterns or your customers. A single API call that ships this week usually beats a training project that takes months.

Training jobs bill only while they run, and batch transform bills only for the job. Real-time endpoints are different: they bill every hour they exist, even with zero traffic. Notebook instances and Studio apps can also keep billing if left running, so always check for leftovers.

A real-time endpoint keeps instances running at all times, so responses are fast and you pay for uptime. Serverless inference provisions compute when a request arrives and scales to zero when idle, so you pay per use but the first request after a quiet period is slower. Pick based on how steady your traffic is.

Any model that serves live decisions for more than a few weeks should be monitored, because drift is silent. A model that is only used for a one-off batch analysis does not need it. Start with data quality monitoring and add model quality monitoring once you can collect ground-truth labels.

No. Bedrock gives you managed access to foundation models for language and reasoning tasks, while SageMaker is for training and serving models on your own structured data. Many production systems use both: SageMaker for a fraud score, Bedrock for summarising the case notes.