What you will learn
- How Docker concepts map directly to ECS — nothing new to learn, just different names
- EC2 Launch Type vs Fargate — what you manage in each and when to choose which
- Task Definitions — the saved configuration for what your container needs to run
- EC2 Instance Profile vs ECS Task Role — two completely separate IAM roles with completely different purposes
- ECS Service Auto Scaling — scaling tasks based on CPU, memory, or request count
- Capacity Providers — proactive cluster scaling that avoids tasks getting stuck
- EFS with ECS for persistent shared storage that survives container restarts
- EventBridge and SQS patterns for event-driven container workloads
- ECR — the AWS container image registry and how it fits into CI/CD
- App Runner — maximum simplicity when you just want a container running
Why this matters
Hotstar runs thousands of containers across dozens of microservices — video metadata, user profiles, recommendations, payments. Each microservice scales independently, recovers from failures automatically, and deploys without downtime. PhonePe runs payment processing microservices on ECS Fargate — no EC2 instances to patch, no cluster capacity to manage, just containers that run and scale automatically.
The difference between a team spending 30% of their time on infrastructure operations versus 5% often comes down to whether containers are configured correctly in ECS.
How Docker Concepts Map to ECS
If you know Docker, ECS adds almost nothing new. Every concept maps directly.
| What you know from Docker | What ECS calls it |
|---|---|
docker run nginx with all its flags |
Task Definition — same config, saved in AWS |
| One running container | Task |
| The machine running containers | Cluster |
docker-compose keeping containers alive |
Service — keeps N tasks always running |
docker push to Docker Hub |
docker push to ECR — same command, different URL |
The full journey:
Write code ↓Write Dockerfile (same as always) ↓docker build (same as always) ↓docker push to ECR (like Docker Hub but on AWS) ↓Task Definition (like your docker run flags, saved permanently) ↓ECS Service (keeps your Tasks running, replaces failed ones) ↓App is liveThe only three new things: ECR (where images live), Task Definition (saved run config), Service (keeps it alive).
ECS Launch Types — EC2 vs Fargate
EC2 Launch Type — you provision the servers:
You create EC2 instances that form your cluster. ECS places containers on those instances for you. But the instances are yours to manage.
Your responsibility with EC2 Launch Type: Provision EC2 instances (choose type, count, AZs) Patch the operating system Manage cluster capacity — if cluster is full, new tasks cannot start Scale EC2 instances up when more capacity is needed ECS responsibility: Decide which container runs on which instance Restart failed containers Register with Load BalancerUse EC2 Launch Type when:
- You need specific instance types (GPU for ML workloads)
- You have Reserved Instances already purchased
- Compliance requires knowing exact server hardware
Fargate — AWS provisions everything:
You say what your container needs (CPU and RAM). AWS figures out where to run it. No EC2 instances in your account. No cluster capacity to manage.
Your responsibility with Fargate: Define CPU and memory in the Task Definition That is it AWS responsibility: Everything else — finding servers, patching them, scaling them To scale with EC2 Launch Type: Add EC2 instances → then scale tasksTo scale with Fargate: Scale tasks directly — AWS handles the rest instantly| EC2 Launch Type | Fargate | |
|---|---|---|
| Infrastructure | You manage EC2 | AWS manages everything |
| Capacity planning | Yes — size your cluster | No — automatic |
| Patching | You patch EC2 OS | None needed |
| Cost model | Pay for EC2 instances | Pay per task second |
| Best for | GPU, Reserved Instances, compliance | Most workloads |
Task Definition — The Recipe for Your Container
A Task Definition is a JSON document that defines everything about how your container should run. Think of it as a saved docker run command.
Task Definition contains: Image URI → which ECR image to pull CPU and Memory → how much compute the task gets Port Mappings → which container port to expose Environment Vars → config the container needs at runtime IAM Role → what AWS services this container can access Log Configuration → where container logs go (CloudWatch) Volumes → EFS mounts for persistent storage Health Check → command to verify container is workingTask Definitions are versioned. Every change creates a new revision. You can roll back to any previous revision instantly.
IAM Roles in ECS — Two Completely Separate Roles
This is the most commonly confused concept in ECS. There are two IAM roles. They serve entirely different purposes. Confusing them breaks security.
EC2 Instance Profile — for the ECS infrastructure:
This role is attached to the EC2 instance (EC2 Launch Type only). The ECS Agent running on the instance uses this role to talk to AWS.
EC2 Instance Profile is used by the ECS Agent to: Register the EC2 instance into the ECS cluster Pull Docker images from ECR Send container logs to CloudWatch Report task status back to ECSThis role is for the plumbing — the infrastructure layer. It has nothing to do with what your application code can access.
ECS Task Role — for your application:
This role is defined inside the Task Definition. Your application code running inside the container uses this role to call AWS services.
Task Role is used by your application code to: Read from S3 Write to DynamoDB Send messages to SQS Call Secrets Manager for credentialsEach task can have its own Task Role with only the permissions it needs. Two tasks on the same EC2 instance can have completely different permissions.
EC2 Instance ├── EC2 Instance Profile → ECS Agent → talks to ECS, ECR, CloudWatch ├── Task A [Task Role A] → reads S3 payments bucket only └── Task B [Task Role B] → writes to DynamoDB orders table onlyRememberEC2 Instance Profile is for the ECS Agent (infrastructure). Task Role is for your application code. Task A cannot access Task B's AWS resources. If you give the EC2 Instance Profile access to S3, ALL tasks on that instance inherit it — this is the wrong approach. Use Task Roles for per-task permissions.
ECS Service Auto Scaling
An ECS Service keeps a desired number of tasks running. Auto Scaling adjusts that desired count based on metrics.
Metrics that trigger task scaling:
ECS Service Average CPU UtilizationECS Service Average Memory UtilizationALB Request Count Per Target (requests per task from ALB)Three strategies (same as ASG):
Target Tracking → pick a target, ECS maintains it automatically Keep average CPU at 60% → ECS adds tasks when above, removes when below Step Scaling → alarm-based tiered responses CPU 70-90% → add 2 tasks CPU above 90% → add 5 tasks Scheduled Scaling → time-based Every weekday 9 AM → set desired to 10 Every weekday 8 PM → set desired to 2Two levels of scaling for EC2 Launch Type:
This trips up a lot of engineers. EC2 Launch Type has TWO independent scaling systems.
ECS Service Auto Scaling → scales the number of TASKS (containers)EC2 Auto Scaling Group → scales the number of EC2 INSTANCES (nodes)They are independent. Adding more tasks does not automatically add more EC2 instances. If the cluster is full, new tasks sit in PENDING state.
Traffic spike: ECS wants to launch 5 more tasks Cluster has no free capacity Tasks stuck in PENDING state — nobody gets served fasterCapacity Providers — smarter than a plain ASG:
Plain ASG: EC2 cluster fills up → tasks go PENDING → ASG eventually notices high EC2 CPU Adds new EC2 → takes 5 minutes → tasks finally place Capacity Provider: ECS wants to place task → no space → Capacity Provider immediately requests new EC2 from ASG before tasks fail → task placed as soon as EC2 is readyCapacity Provider tells ECS and ASG to work together proactively instead of each reacting independently.
TipFor EC2 Launch Type, always use a Capacity Provider Strategy instead of managing the ASG separately. It prevents tasks from getting stuck in PENDING state during scale-out events — the most frustrating ECS operational problem.
ECS with EFS — Persistent Shared Storage
Containers are stateless by design. Data written inside a container is lost when the container stops. For data that must survive — uploaded files, shared configuration, ML model files — mount Amazon EFS.
Fargate Task A in ap-south-1a Fargate Task B in ap-south-1b /uploads (EFS mounted) /uploads (EFS mounted) | | └───────────────────────────────┘ | Amazon EFS (data lives here, shared)Both tasks read and write to the same EFS file system. Data persists when containers stop and restart. New tasks mounting the same EFS see all existing files immediately.
EFS works with both EC2 and Fargate launch types. It is the only shared persistent storage option for Fargate tasks — EBS cannot attach to Fargate.
ECS Integration Patterns
ECS with EventBridge — run tasks on events:
S3 object uploaded (new video for transcoding) ↓EventBridge rule: source=s3, event=ObjectCreated ↓ECS Task launched automatically (Fargate) ↓Task transcodes the video → saves output → stopsPay only for the time the task ranNo polling. No always-on workers. Tasks start on demand when events arrive.
ECS with SQS — queue-based processing:
Orders arrive in SQS queue ↓ECS Service polls queue ↓Queue depth grows → CloudWatch metric rises ↓ECS Service Auto Scaling adds more tasks ↓More tasks = faster processing = queue drainsIntercepting stopped tasks with EventBridge:
ECS Task stops unexpectedly (crash, OOM, timeout) ↓ECS emits a Task State Change event to EventBridge ↓EventBridge rule matches: stoppedReason contains "OutOfMemoryError" ↓SNS alert sent to ops team immediatelyECR — Container Image Registry
ECR (Elastic Container Registry) is AWS's managed Docker image registry. Like Docker Hub but hosted inside AWS with IAM integration.
Push workflow:
## Step 1 — Authenticate Docker to ECR (token valid 12 hours)aws ecr get-login-password --region ap-south-1 | \ docker login --username AWS --password-stdin \ 123456789012.dkr.ecr.ap-south-1.amazonaws.com ## Step 2 — Create the repository (one-time)aws ecr create-repository \ --repository-name devops-orders-api \ --image-scanning-configuration scanOnPush=true \ --region ap-south-1 ## Step 3 — Build, tag, and pushdocker build -t devops-orders-api:v1 .docker tag devops-orders-api:v1 \ 123456789012.dkr.ecr.ap-south-1.amazonaws.com/devops-orders-api:v1docker push \ 123456789012.dkr.ecr.ap-south-1.amazonaws.com/devops-orders-api:v1ECR Lifecycle Policy — auto-delete old images:
Without lifecycle policies, every CI/CD push stores a new image permanently. Over months this becomes thousands of images costing significant money.
ECR → Repositories → devops-orders-api → Lifecycle Policy → Create ruleRule priority: 1Image status: AnyMatch criteria: Image count more than 10Action: Expire → SaveThis keeps only the 10 most recent images and deletes older ones automatically.
Common MistakeIf ECS cannot pull an image from ECR, check IAM first. The most common cause is a missing policy on the EC2 Instance Profile (for EC2 Launch Type) or the Task Execution Role (for Fargate). The execution role needs: ecr:GetAuthorizationToken, ecr:BatchCheckLayerAvailability, and ecr:GetDownloadUrlForLayer.
App Runner — Maximum Simplicity
App Runner is the easiest way to run a containerised app on AWS. Give it your container image or source code — App Runner builds, deploys, runs it, and handles HTTPS, load balancing, and auto-scaling automatically.
Container image or GitHub repository ↓Configure: vCPU, RAM, auto-scaling range, health check ↓Deploy ↓HTTPS URL → your app is live App Runner handles: Building from source (if using GitHub) Deploying automatically on every push Auto-scaling up and down Load balancing HTTPS certificates ECS/EKS → full control, more configuration, more operational responsibilityApp Runner → maximum simplicity, no infrastructure decisions at allUse App Runner for: web apps, APIs, internal tools, microservices where you want speed over control.
Hands-on Lab — Deploy Container on Fargate with ALB
Step 1 — Create an ECS cluster
ECS → Clusters → Create clusterCluster name: devops-prod-clusterInfrastructure: AWS Fargate (serverless) ← tick thisCreateStep 2 — Create a Task Definition
ECS → Task definitions → Create new task definitionTask definition family: devops-orders-apiInfrastructure requirements: Launch type: Fargate CPU: 0.5 vCPU Memory: 1 GB Container: Name: orders-api Image URI: nginx:latest (use your ECR image in production) Port mappings: Container port 80, Protocol HTTP Log collection: Use log collection → creates CloudWatch log group automatically CreateStep 3 — Create an ECS Service
ECS → devops-prod-cluster → Services → Create Deployment configuration: Application type: Service Task definition: devops-orders-api (LATEST) Service name: orders-api-service Desired tasks: 2 Networking: VPC: your VPC Subnets: select your private subnets Security Group: create new → allow port 80 from ALB SG Load balancing: Load balancer type: Application Load Balancer Create a new ALB: devops-orders-alb Listener: HTTP 80 Target group: create new → health check path: / Create Service creation takes 3-5 minutes. Both tasks should reach RUNNING state.Step 4 — Verify tasks are running
ECS → devops-prod-cluster → Services → orders-api-serviceTasks tab → both tasks should show Status: RUNNING Click on one task → Logs tab → you should see nginx startup logs in CloudWatchStep 5 — Test through the ALB
EC2 → Load Balancers → devops-orders-alb → DNS nameOpen in browser: http://devops-orders-alb-123456.ap-south-1.elb.amazonaws.comYou should see the nginx welcome page served by your Fargate taskStep 6 — Add Auto Scaling to the service
ECS → orders-api-service → Update serviceService Auto Scaling → EditMinimum tasks: 2Maximum tasks: 10 Add scaling policy: Policy name: cpu-target-60 Policy type: Target tracking ECS service metric: ECSServiceAverageCPUUtilization Target value: 60% Update serviceStep 7 — Cleanup
ECS → orders-api-service → Delete service (drains tasks first)ECS → devops-prod-cluster → Delete clusterEC2 → Load Balancers → devops-orders-alb → DeleteEC2 → Target Groups → delete the target groupCommon Mistakes to Avoid
Common MistakeUsing the EC2 Instance Profile to give containers access to AWS services. If the Instance Profile has S3 access, every container on that EC2 can access S3 — including containers that should not have that permission. Use Task Roles. Each task gets exactly the permissions it needs and nothing more.
Common MistakeNot configuring a Capacity Provider for EC2 Launch Type and wondering why tasks get stuck in PENDING during scale-out. Without a Capacity Provider, ECS and the ASG operate independently. The ASG does not know tasks are waiting. Tasks sit PENDING until the ASG eventually scales for unrelated reasons. Use Capacity Providers from day one.
Common MistakeForgetting to set up ECR Lifecycle Policies and discovering 6 months later that storage costs are significant due to thousands of accumulated images. Every CI/CD push creates a new image. Without lifecycle policies, none of them are ever deleted. Set it up when you create the repository.
TipWhen a Fargate task stops immediately after starting, always check CloudWatch Logs first. Fargate sends stdout and stderr to CloudWatch if you configure the awslogs log driver (which the console does automatically). The exit reason in CloudWatch almost always tells you exactly what went wrong in one line — OOM error, app crash, missing environment variable, failed health check command.