What you will learn
- When to choose EKS over ECS — the open standard vs AWS-native decision
- EKS architecture — what AWS manages vs what you manage
- The three node types — Managed Node Groups, Self-Managed Nodes, and Fargate Profiles
- EKS networking — the VPC CNI plugin and why every pod gets a real VPC IP
- EKS storage — EBS CSI driver for one pod, EFS CSI driver for many pods
- IRSA — IAM Roles for Service Accounts — the correct way to give pods AWS permissions
- Why IRSA is better than EC2 Instance Profile for pod-level permissions
- EKS add-ons managed by AWS — CoreDNS, kube-proxy, VPC CNI, EBS CSI
- AWS Load Balancer Controller for Ingress and LoadBalancer Service resources
- How to connect to your EKS cluster with kubectl after creation
Why this matters
PhonePe runs their payment microservices on EKS — the Kubernetes ecosystem gives their platform team portability across clouds, access to the full CNCF tooling stack, and a consistent deployment model whether running on AWS, on-premises, or in any other cloud. The Kubernetes API is the same everywhere. Teams that have invested in Kubernetes expertise — Helm charts, GitOps with ArgoCD, Prometheus monitoring, Istio service mesh — can bring all of that to EKS without changes. ECS is simpler and AWS-native. EKS is the right choice when the Kubernetes ecosystem investment is already made or the organisation needs cloud portability.
ECS vs EKS — When to Choose Which
ECS: AWS-proprietary API and tooling Simpler learning curve — fewer concepts Best for: AWS-first teams, teams new to containers Fargate support: yes Cloud portability: AWS only Kubernetes ecosystem (Helm, ArgoCD, Istio): no EKS: Kubernetes open standard — same API as GKE, AKS, on-premises Higher learning curve — full Kubernetes concepts required Best for: teams already using Kubernetes, multi-cloud strategy Fargate support: yes (via Fargate Profiles) Cloud portability: move to any cloud or on-premises Kubernetes ecosystem: full access to all CNCF tools Decision guide: Your team already knows Kubernetes → EKS You are migrating on-premises K8s to AWS → EKS Multi-cloud or cloud-agnostic is a priority → EKS You are starting fresh with no K8s knowledge → ECS You want maximum AWS-native simplicity → ECSRememberECS is mentioned in questions as "AWS containers". EKS is mentioned as "Kubernetes". When a question uses the word Kubernetes, EKS is the answer. When it describes a simpler AWS-native container solution, ECS is the answer.
EKS Architecture
EKS splits responsibility between AWS and you.
AWS manages (the Control Plane): Kubernetes API server etcd (cluster state database) Controller Manager Scheduler Control Plane runs across multiple AZs automatically AWS guarantees 99.95% API server uptime You manage (the Worker Nodes — unless using Fargate): EC2 instances running your pods Node OS patching (partially with Managed Node Groups) Storage attached to nodes Networking configuration Architecture overview: AWS Cloud (VPC) ├── AZ-a: Public subnet (ALB, NAT GW) | Private subnet (EKS Node → Pods) ├── AZ-b: Public subnet (ALB, NAT GW) | Private subnet (EKS Node → Pods) └── AZ-c: Public subnet (ALB, NAT GW) | Private subnet (EKS Node → Pods) | Auto Scaling Group (manages adding and removing nodes)Nodes run in private subnets — never exposed to the internet directly. Load Balancers sit in public subnets and route traffic inward. Nodes spread across multiple AZs for high availability.
EKS Key Terminology
| EKS Term | ECS Equivalent | What It Is |
|---|---|---|
| Pod | Task | Smallest unit — one or more containers running together |
| Node | EC2 Instance | The server that runs your pods |
| Node Group | EC2 Launch Type | A group of EC2 instances managed by an ASG |
| Fargate Profile | Fargate Launch Type | Serverless — no nodes to manage |
| Deployment | Service | Keeps a set number of pod replicas running |
| Service | Target Group + ALB | DNS-based load balancing to a set of pods |
| Ingress | ALB + routing rules | HTTP routing rules into the cluster |
| Namespace | — | Virtual cluster for resource isolation |
EKS Node Types
Managed Node Groups:
AWS creates and manages the underlying EC2 instances and their ASG. You choose the instance type, desired size, and scaling bounds. AWS handles OS patching for the node AMI via rolling updates.
You control: instance type, size limits, labels, taintsAWS controls: underlying EC2 provisioning, ASG, node AMI updates Supports Spot Instances — Spot node groups for cost savingsNodes can be in multiple AZs for HABest for: standard production workloads where you want managed nodesSelf-Managed Nodes:
You create your own EC2 instances and register them into the EKS cluster manually using the bootstrap script. Full control — custom AMIs, custom instance configurations, any EC2 type.
You control: everything — EC2, patching, ASG, custom AMIsBest for: highly specific hardware requirements, custom OS configurationsAWS Fargate Profiles:
Pods run serverlessly. No nodes to manage. You define a Fargate Profile specifying which namespace and pod selectors run on Fargate. Pods matching the selector are automatically scheduled on Fargate.
You define: namespace and label selectors for which pods use FargateAWS handles: the underlying compute — invisible to youLimitation: Fargate pods can only use EFS storage, not EBS| Managed Node Groups | Self-Managed Nodes | Fargate | |
|---|---|---|---|
| Who manages EC2 | AWS creates and manages | You create and manage | No EC2 at all |
| Control | Medium | Full | None needed |
| Maintenance | Low | High | Zero |
| Spot support | Yes | Yes | No |
| EBS storage | Yes | Yes | No |
| EFS storage | Yes | Yes | Yes |
EKS Networking — VPC CNI Plugin
EKS uses the AWS VPC CNI plugin for pod networking. This makes EKS networking unique compared to most Kubernetes distributions.
Standard Kubernetes networking: Pods get IP addresses from a separate pod CIDR range (overlay network) Pod-to-pod traffic goes through a virtual network overlay External traffic must be translated from pod IP to node IP (NAT) EKS with VPC CNI: Every pod gets a REAL VPC IP address from your subnet Pods are visible as actual ENI secondary IPs on the node Pod-to-pod traffic routes directly via VPC — no overlay Security Groups can be applied directly to pods No NAT required — pods are first-class VPC citizensThis means:
Pod IP 10.0.1.47 → reachable directly from any resource in the VPCRDS can allow inbound from a Pod Security Group ID directlyNo translation layer between pod and AWS networkingRememberWith VPC CNI, pods consume real VPC IP addresses from your subnets. Large clusters with many pods can exhaust subnet IP space quickly. Plan your subnet CIDR ranges to be large enough — /20 or larger — when designing EKS subnets.
EKS Storage — CSI Drivers
To attach storage to EKS pods you install CSI (Container Storage Interface) drivers and define StorageClass objects.
Amazon EBS CSI Driver:
Provides block storage (like a hard drive) to podsOne EBS volume → one pod at a time (ReadWriteOnce)Works only with EC2 node groups — not Fargate podsUse for: databases, stateful applications needing dedicated block storage## StorageClass using EBS gp3apiVersion: storage.k8s.io/v1kind: StorageClassmetadata: name: ebs-gp3provisioner: ebs.csi.aws.comparameters: type: gp3 encrypted: "true"volumeBindingMode: WaitForFirstConsumerAmazon EFS CSI Driver:
Provides shared file storage across multiple pods simultaneously (ReadWriteMany)Works with BOTH EC2 node groups AND Fargate podsData persists beyond pod lifetimeUse for: shared content, multi-pod write access, Fargate persistent storage## StorageClass using EFSapiVersion: storage.k8s.io/v1kind: StorageClassmetadata: name: efs-scprovisioner: efs.csi.aws.comparameters: provisioningMode: efs-ap fileSystemId: fs-0abc123456789def directoryPerClaim: "true"RememberEFS is the only storage option that works with both EC2 nodes AND Fargate pods. EBS only works with EC2 nodes. If a pod on Fargate needs persistent storage, the answer is always EFS.
| Storage | Access Mode | Works with Fargate | Use Case |
|---|---|---|---|
| EBS | ReadWriteOnce (one pod) | No | Databases, dedicated block storage |
| EFS | ReadWriteMany (many pods) | Yes | Shared files, multi-pod access |
| FSx for Lustre | ReadWriteMany | No | HPC, ML training data |
| FSx for NetApp ONTAP | ReadWriteMany | No | Enterprise NAS workloads |
IRSA — IAM Roles for Service Accounts
IRSA is the correct way to give pods permission to access AWS services. Without IRSA, pods on EC2 nodes inherit the node's IAM Instance Profile — which means ALL pods on a node get the same permissions. This violates least privilege.
The problem with EC2 Instance Profile for pods:
Node has Instance Profile: S3FullAccess + DynamoDBFullAccess + SQSFullAccess Pod A (orders service) → needs only SQS and DynamoDB accessPod B (media service) → needs only S3 accessPod C (reporting) → needs only DynamoDB read access Without IRSA: ALL three pods get S3Full + DynamoDB Full + SQSFullWith IRSA: each pod gets exactly and only what it needsHow IRSA works:
1. Create an IAM Role with the permissions the pod needs2. Add a trust policy allowing the EKS OIDC provider to assume the role3. Annotate the Kubernetes Service Account with the IAM Role ARN4. Deploy the pod using that Service Account5. AWS SDK in the pod automatically fetches temporary credentials for that role## Step 1 — Get the OIDC issuer URL for your clusteraws eks describe-cluster \ --name devops-prod-cluster \ --query "cluster.identity.oidc.issuer" \ --output text \ --region ap-south-1## Output: https://oidc.eks.ap-south-1.amazonaws.com/id/EXAMPLED539D4633E53DE1B71EXAMPLE ## Step 2 — Create IAM OIDC provider for the clusteraws iam create-open-id-connect-provider \ --url https://oidc.eks.ap-south-1.amazonaws.com/id/EXAMPLED539D4633E53DE1B71EXAMPLE \ --client-id-list sts.amazonaws.com \ --thumbprint-list 9e99a48a9960b14926bb7f3b02e22da2b0ab7280 ## Step 3 — Create IAM role with trust policy for the Service Accountcat > trust-policy.json << 'EOF'{ "Version": "2012-10-17", "Statement": [{ "Effect": "Allow", "Principal": { "Federated": "arn:aws:iam::123456789012:oidc-provider/oidc.eks.ap-south-1.amazonaws.com/id/EXAMPLED539D4633E53DE1B71EXAMPLE" }, "Action": "sts:AssumeRoleWithWebIdentity", "Condition": { "StringEquals": { "oidc.eks.ap-south-1.amazonaws.com/id/EXAMPLED539D4633E53DE1B71EXAMPLE:sub": "system:serviceaccount:devops-namespace:orders-service-account" } } }]}EOF aws iam create-role \ --role-name devops-orders-pod-role \ --assume-role-policy-document file://trust-policy.json ## Attach only what this pod needsaws iam attach-role-policy \ --role-name devops-orders-pod-role \ --policy-arn arn:aws:iam::aws:policy/AmazonSQSFullAccess## Step 4 — Annotate the Kubernetes Service AccountapiVersion: v1kind: ServiceAccountmetadata: name: orders-service-account namespace: devops-namespace annotations: eks.amazonaws.com/role-arn: arn:aws:iam::123456789012:role/devops-orders-pod-role## Step 5 — Deploy pod using the annotated Service AccountapiVersion: apps/v1kind: Deploymentmetadata: name: orders-api namespace: devops-namespacespec: replicas: 2 selector: matchLabels: app: orders-api template: metadata: labels: app: orders-api spec: serviceAccountName: orders-service-account ## ← uses our annotated SA containers: - name: orders-api image: 123456789012.dkr.ecr.ap-south-1.amazonaws.com/orders-api:latest ports: - containerPort: 3000The AWS SDK in the pod automatically receives temporary credentials for devops-orders-pod-role. No access keys anywhere. No node-level permissions shared across pods.
AWS Load Balancer Controller
The AWS Load Balancer Controller is a Kubernetes controller that watches for Ingress and Service objects and creates ALBs and NLBs automatically.
You create a Kubernetes Ingress object ↓AWS Load Balancer Controller detects it ↓Creates an ALB in your VPC automatically ↓Configures routing rules, target groups, health checks ↓Users hit the ALB → routed to the correct pods## Ingress object that creates an ALB automaticallyapiVersion: networking.k8s.io/v1kind: Ingressmetadata: name: devops-ingress namespace: devops-namespace annotations: kubernetes.io/ingress.class: alb alb.ingress.kubernetes.io/scheme: internet-facing alb.ingress.kubernetes.io/target-type: ip alb.ingress.kubernetes.io/certificate-arn: arn:aws:acm:ap-south-1:123456789012:certificate/abc123spec: rules: - host: api.devops-network.in http: paths: - path: /orders pathType: Prefix backend: service: name: orders-service port: number: 80 - path: /users pathType: Prefix backend: service: name: users-service port: number: 80EKS Add-ons Managed by AWS
EKS Add-ons are operational components that AWS manages — you do not patch or upgrade them manually.
| Add-on | What it does |
|---|---|
| CoreDNS | DNS resolution inside the cluster — pods resolve service names |
| kube-proxy | Network rules on each node for Service routing |
| Amazon VPC CNI | Pod networking — gives every pod a real VPC IP |
| Amazon EBS CSI Driver | Block storage provisioning for pods |
| Amazon EFS CSI Driver | Shared file storage provisioning |
| AWS Load Balancer Controller | Creates ALB and NLB from Kubernetes objects |
## List installed add-onsaws eks list-addons \ --cluster-name devops-prod-cluster \ --region ap-south-1 ## Get add-on details including versionaws eks describe-addon \ --cluster-name devops-prod-cluster \ --addon-name vpc-cni \ --region ap-south-1 ## Update an add-on to latest versionaws eks update-addon \ --cluster-name devops-prod-cluster \ --addon-name vpc-cni \ --resolve-conflicts OVERWRITE \ --region ap-south-1Hands-on Lab — Create EKS Cluster and Deploy Application
Step 1 — Install eksctl and create the cluster
eksctl is the official EKS CLI — it handles cluster creation, IAM roles, VPC, and OIDC provider all in one command. Download from eksctl.io.
## Create a complete EKS cluster in one commandeksctl create cluster --name devops-prod-cluster --region ap-south-1 --nodegroup-name devops-nodes --node-type t3.medium --nodes 2 --nodes-min 2 --nodes-max 5## Takes 15-20 minutes. Creates VPC, subnets, IAM roles, OIDC provider automatically.Step 2 — Connect kubectl to your cluster
aws eks update-kubeconfig --region ap-south-1 --name devops-prod-clusterkubectl get nodes## Both nodes should show: ReadyStep 3 — Deploy nginx
kubectl create deployment devops-api --image=nginx --replicas=2kubectl expose deployment devops-api --port=80 --type=ClusterIPkubectl get pods ## both pods Runningkubectl get svc ## service createdStep 4 — Set up IRSA so pods access S3 safely
Instead of giving the whole EC2 node S3 access (wrong approach),IRSA gives only specific pods the permissions they need.## Create a service account linked to an IAM role — one commandeksctl create iamserviceaccount --name s3-reader-sa --namespace default --cluster devops-prod-cluster --attach-policy-arn arn:aws:iam::aws:policy/AmazonS3ReadOnlyAccess --approve --region ap-south-1 ## Verify — the annotation is the link between K8s SA and IAM rolekubectl describe serviceaccount s3-reader-sa## Look for: eks.amazonaws.com/role-arn annotationStep 5 — Inspect and scale
kubectl scale deployment devops-api --replicas=4kubectl get pods --watch ## watch 2 more pods appear kubectl logs -l app=devops-api ## see nginx logskubectl describe pod $(kubectl get pods -o name | head -1) ## full pod detailsStep 6 — Cleanup
kubectl delete deployment devops-apikubectl delete svc devops-apieksctl delete cluster --name devops-prod-cluster --region ap-south-1## Deletes everything — takes 10-15 minutesProduction Best Practices and Common Pitfalls
- Use IRSA for all pod-level AWS permissions — never rely on the node Instance Profile for pod permissions
- Plan subnet CIDR ranges carefully — VPC CNI assigns real VPC IPs to pods, large clusters exhaust subnets quickly
- Use Managed Node Groups for standard workloads — self-managed nodes add operational overhead without benefit in most cases
- Install the AWS Load Balancer Controller for Ingress resources — do not use the deprecated in-tree cloud controller
- Enable OIDC provider for the cluster immediately after creation — IRSA depends on it and retrofitting it later is an extra step
- Use EFS for any storage that Fargate pods need — EBS cannot attach to Fargate
- Enable cluster logging (API server, audit, authenticator, controller manager, scheduler) — essential for debugging and compliance
Quick Reference and Troubleshooting Commands
| Task | Command |
|---|---|
| Create cluster | aws eks create-cluster --name <name> --kubernetes-version 1.30 --role-arn <arn> --resources-vpc-config subnetIds=<ids> |
| Update kubeconfig | aws eks update-kubeconfig --region ap-south-1 --name <cluster-name> |
| List clusters | aws eks list-clusters --region ap-south-1 |
| List node groups | aws eks list-nodegroups --cluster-name <name> --region ap-south-1 |
| Get nodes | kubectl get nodes -o wide |
| Get all pods | kubectl get pods --all-namespaces |
| Describe pod | kubectl describe pod <pod-name> |
| Get pod logs | kubectl logs <pod-name> --follow |
| Exec into pod | kubectl exec -it <pod-name> -- /bin/bash |
| Scale deployment | kubectl scale deployment <name> --replicas=5 |
| Apply manifest | kubectl apply -f manifest.yaml |
| Delete resources | kubectl delete -f manifest.yaml |
Common problems and fixes:
| Problem | Likely cause | Fix |
|---|---|---|
| Pods stuck in Pending | No nodes with available capacity | Check node capacity, add nodes or enable Cluster Autoscaler |
| Pod cannot reach AWS service | IRSA not configured or wrong role ARN | Verify Service Account annotation, check OIDC provider created |
| Pods exhausting subnet IPs | Subnet too small for pod count | Use /20 or larger subnets, or enable prefix delegation on VPC CNI |
| Node NotReady | Node failed health check | Check node logs: kubectl describe node |
| ImagePullBackOff | ECR auth expired or wrong registry URL | Check execution role has ECR permissions, verify image URI |
Common MistakeUsing the node EC2 Instance Profile to give pods access to AWS services. All pods on that node get the same permissions — a compromised pod can access everything the node can access. Use IRSA. Each pod gets exactly the permissions it needs and nothing more.
TipUse eksctl (the official EKS CLI tool) instead of raw AWS CLI for cluster creation and management. A single eksctl create cluster command sets up the VPC, subnets, node groups, IAM roles, and OIDC provider automatically. What takes 20 CLI commands with aws eks takes one command with eksctl. Install from eksctl.io.
Common Mistakes to Avoid
Common MistakeUsing the node EC2 Instance Profile to give pods access to AWS services. All pods on that node inherit the same permissions — a compromised pod can access everything the node can. Use IRSA. Each pod gets exactly the permissions it needs and nothing more.
Common MistakeUsing small subnets for EKS. VPC CNI assigns real VPC IPs to every pod. A /27 subnet with 27 usable IPs fills up fast on a busy cluster. Use /20 or larger subnets for EKS workloads.
TipUse eksctl instead of raw AWS CLI for cluster management. What takes 15 CLI commands takes one eksctl command. It also handles the OIDC provider setup that IRSA depends on — which many engineers forget when using raw CLI.