What you will learn
- What EC2 is and why it is the foundation every other AWS service builds on
- AMI types — AWS-provided, Marketplace, and custom — and what each contains
- The EC2 instance family naming format and which family to choose for which workload
- EC2 User Data — the bootstrap script that configures servers automatically at first launch
- Security Groups — the virtual firewall, inbound/outbound rules, and chaining patterns
- Key Pairs and SSH — four methods to connect to EC2 instances
- Public IP vs Private IP vs Elastic IP and what changes on stop/start
- IAM Roles — why they are always better than access keys on EC2
- All EC2 purchasing options — On-Demand, Reserved, Spot, Savings Plans, Dedicated Hosts
- Placement Groups — Cluster, Spread, Partition and the production use case for each
- Elastic Network Interfaces — the failover pattern that keeps IP addresses stable
- EC2 Hibernate — when and why to use it over a normal stop
Why this matters
Every production workload at Swiggy, Razorpay, Zerodha, and Hotstar runs on compute. EC2 is that compute for most teams. Understanding EC2 deeply is not optional — it is the foundation. A wrong instance type means money wasted or performance problems. A poorly configured Security Group means a security incident. Not using IAM Roles and using access keys instead is one of the most common AWS security breaches. Getting EC2 right from day one is what separates infrastructure that scales safely from infrastructure that eventually becomes a liability.
What is EC2
EC2 (Elastic Compute Cloud) is AWS's service for renting virtual servers in the cloud. Instead of buying physical hardware, you rent compute capacity from AWS and pay only for what you use.
EC2 is Infrastructure as a Service (IaaS) — AWS manages the physical hardware, you manage everything above it: the OS, software, security, and applications.
EC2 gives you: Renting virtual machines → EC2 instances Storing data on virtual drives → EBS volumes Distributing load across machines → Elastic Load Balancer Scaling services automatically → Auto Scaling GroupsCommon use cases:
Web servers and APIsBackend application serversDocker and Kubernetes workloads (nodes)Batch processing and data jobsMachine learning training and inferenceCI/CD build agentsAMI — Amazon Machine Image
An AMI is a pre-built template containing everything needed to launch an EC2 instance — the operating system, pre-installed software, and configuration. Every EC2 instance starts from an AMI.
AMIs are built for a specific AWS region and can be copied across regions.
Three types of AMIs:
AWS-provided → Amazon Linux 2, Ubuntu, Windows Server — maintained by AWSAWS Marketplace → vendor-built with software pre-installed (WordPress, LAMP stack, etc)Custom AMI → you build from a running instance, pre-bake your appCommon AMIs and default usernames:
| AMI | Default Username | Common use |
|---|---|---|
| Amazon Linux 2 | ec2-user | General purpose, AWS-optimized |
| Ubuntu 22.04 | ubuntu | Web servers, DevOps workloads |
| Windows Server | Administrator | Windows-based applications |
| Red Hat | ec2-user | Enterprise Linux workloads |
| Debian | admin | Stable Linux environments |
AMI and EBS snapshots:
When you create a custom AMI from a running instance, AWS automatically snapshots every attached EBS volume.
You click Create AMI on a running instance ↓AWS snapshots all attached EBS volumes automatically ↓AMI stores references to those snapshots (not the actual data) ↓You launch a new EC2 from that AMI ↓AWS uses those snapshots to create fresh EBS volumes ↓New EC2 boots with the exact same disk state as the originalRememberDeleting the AMI does NOT automatically delete the underlying snapshots. Orphaned snapshots keep billing you. Always clean them up manually after deleting an AMI.
Instance Types — Choosing the Right Size
The instance type defines CPU, RAM, network speed, and storage performance. Naming format:
m5.2xlarge| || +-- size: micro, small, medium, large, xlarge, 2xlarge, 4xlarge...+-- family: m = general purpose, c = compute, r = memory...The number in the family name is the generation — higher is newer and better. m5 is newer than m4.
General Purpose (T, M families) — balanced CPU, memory, network:
Web servers and APIsSmall to medium databasesDevelopment and test environmentsMicroservices Examples: t2.micro (Free Tier), t3.medium, m5.large, m6i.xlargeCompute Optimized (C family) — high-performance processors:
Batch processing workloadsMedia transcodingHigh Performance ComputingScientific modelingDedicated gaming servers Examples: c5.large, c6i.xlarge, c6g.2xlargeMemory Optimized (R, X families) — large datasets in RAM:
High-performance relational and non-relational databasesIn-memory databases (Redis at very large scale)Real-time processing of large unstructured data Examples: r5.large, r6g.xlarge, x1e.32xlargeStorage Optimized (I, D families) — high sequential read/write:
High-frequency OLTP systemsRelational and NoSQL databasesDistributed file systems (HDFS)Data warehousing Examples: i3.large, i4i.xlarge, d2.xlargeAccelerated Computing (P, G families) — GPU workloads:
Machine learning training and inferenceGraphics rendering and video encoding Examples: p3.2xlarge, g4dn.xlargeTipCompare all instance types side by side at https://instances.vantage.sh — shows pricing, vCPU, RAM, network, and storage for every instance type in every region.
EC2 User Data — Automated Bootstrap
EC2 User Data is a script that runs automatically when an EC2 instance starts for the very first time. It automates initial server setup so you never SSH in and configure manually.
Key rules:
Runs ONLY once — at the very first boot, never again afterRuns as the root user — no sudo needed inside the scriptMust start with: #!/bin/bashIdeal for: install software, update packages, start services, download files## Update all installed packagesyum update -y ## Install nginx web serveryum install -y nginx ## Start nginx immediatelysystemctl start nginx ## Enable nginx to auto-start on every future rebootsystemctl enable nginx ## Create homepage to verify server is workingecho "<h1>Server deployed via User Data on $(hostname -f)</h1>" > /var/www/html/index.htmlWithout User Data: Launch EC2 → SSH in → run updates → install nginx → configure → repeat for every server With User Data: Launch EC2 with script → server ready when it boots Same script for 1 server or 1000 — identical every timeRememberUser Data only runs at first launch. If you need to re-run it, you must stop the instance, modify the User Data, and start again. For ongoing configuration management use AWS Systems Manager or Ansible.
Security Groups
A Security Group is a virtual firewall controlling all inbound and outbound traffic to and from your EC2 instance. Every EC2 instance must have at least one Security Group attached.
Three key characteristics:
1. Allow rules only — no explicit deny:
Rule: Allow port 80 from 0.0.0.0/0Result: HTTP traffic allowed, everything else blocked automatically2. Stateful — return traffic automatically allowed:
User sends HTTP request to port 80 ↓Inbound rule: Allow port 80 → CHECKED → request reaches EC2 ↓EC2 sends response back ↓Response automatically allowed — NO separate outbound rule needed3. Region and VPC specific:
A Security Group in ap-south-1 cannot attach to an instance in us-east-1Common ports to know:
| Port | Protocol | Purpose |
|---|---|---|
| 22 | SSH | Remote login to Linux |
| 80 | HTTP | Unsecured web traffic |
| 443 | HTTPS | Secured web traffic |
| 3389 | RDP | Remote desktop for Windows |
| 3306 | MySQL/Aurora | Database access |
| 5432 | PostgreSQL | Database access |
Security Group chaining — the three-tier production pattern:
Internet ↓ allow 80, 443 from 0.0.0.0/0Web Server SG ↓ allow 8080 from Web Server SG onlyApp Server SG ↓ allow 5432 from App Server SG onlyDatabase SG → RDSThe database Security Group source is the App Server Security Group ID — not an IP address. IPs change. Security Group IDs never change.
SecurityIf your instance shows a connection timeout error, it is almost always a Security Group issue — a port is not open. If it shows connection refused, the Security Group is fine but the application is not running on that port. These two errors tell you exactly where the problem is.
Key Pairs and SSH Access
A Key Pair uses asymmetric cryptography instead of passwords.
Public key → AWS stores this on the EC2 instance during launchPrivate key → you download the .pem file and keep it on your machineMethod 1 — SSH on Linux and macOS:
## Set correct permissions — SSH refuses keys that are too openchmod 400 your-key.pem ## Connect to the instancessh -i your-key.pem ubuntu@13.235.45.67 ## Command breakdown:## -i your-key.pem → specify the identity (key) file## ubuntu → username (depends on the AMI)## 13.235.45.67 → EC2 public IP addressMethod 2 — SSH on Windows (PowerShell or Git Bash):
ssh -i your-key.pem ec2-user@13.235.45.67## Windows 10 and 11 have SSH built inMethod 3 — EC2 Instance Connect (browser-based, no key needed):
EC2 Dashboard → Select Instance → Connect → EC2 Instance Connect → ConnectAWS uploads a temporary one-time key automatically. Port 22 must still be open in the Security Group.
Method 4 — SSM Session Manager (no port 22, no key needed):
EC2 Dashboard → Connect → Session Manager → ConnectNo inbound ports required. Instance must have the SSM Agent installed and an IAM Role with AmazonSSMManagedInstanceCore policy. The most secure option for production.
IP Addressing
Public IP:
Accessible from the internetGlobally uniqueAssigned when instance launchesCHANGES every time you stop and start the instance (unless using Elastic IP)Used for SSH: ssh ec2-user@16.16.123.7Private IP:
Accessible only within the VPCNever changes — even after stop and startUsed for internal communication between serversElastic IP:
A static public IPv4 address you own and assign to an EC2 instance.
Stays the same even after instance stop and startFree when attached to a running instanceBILLED when not attached — AWS charges for unused Elastic IPsMaximum 5 per account by default Failover use case: Users → Elastic IP 54.32.11.22 → Instance A (running) Instance A fails → remap Elastic IP → Instance B (backup) Users still hit 54.32.11.22 → now reaches Instance B Zero DNS changes neededTipAvoid Elastic IPs where possible. A better approach is using a Load Balancer with a stable DNS name so you never depend on a raw IP address. Elastic IPs reflect a design that should probably use a load balancer instead.
IAM Roles with EC2
If your EC2 instance needs to call AWS services (S3, DynamoDB, Secrets Manager), the naive approach is running aws configure on the instance and entering personal access keys.
Keys stored on the instance → instance compromised → your credentials exposedAttacker can do anything your account allowsThe correct approach — IAM Roles:
Create IAM Role (e.g. EC2-S3-ReadOnly-Role) ↓Attach a Policy (e.g. AmazonS3ReadOnlyAccess) ↓Attach the Role to the EC2 instance (IAM Instance Profile) ↓EC2 automatically receives TEMPORARY, ROTATING credentials ↓No keys stored anywhere. Nothing to leak. Nothing to rotate manually.## On EC2 with no role attached:aws s3 ls## Error: Unable to locate credentials ## On EC2 with S3 read role attached:aws s3 ls## Works automatically — no aws configure neededAttach role to running instance:
EC2 Dashboard → Select Instance → Actions → Security → Modify IAM Role → Select Role → UpdateSecurityNever store AWS access keys on an EC2 instance. This is one of the most common and dangerous mistakes in cloud security. Use IAM Roles. The SDK fetches temporary credentials automatically. No keys. No risk.
EC2 Purchasing Options
| Option | Savings | Commitment | Best For |
|---|---|---|---|
| On-Demand | None (baseline) | None | Unpredictable workloads |
| Reserved Instances | Up to 72% | 1 or 3 years | Steady 24/7 workloads |
| Savings Plans | Up to 72% | 1 or 3 years | Flexible long-term usage |
| Spot Instances | Up to 90% | None | Fault-tolerant batch jobs |
| Dedicated Hosts | None | Variable | BYOL and compliance |
| Capacity Reservations | None | None | Guaranteed AZ capacity |
On-Demand: Pay by hour or second. No commitment. Highest per-hour rate. Start/stop anytime.
Reserved Instances: Commit to specific instance type, region, OS for 1 or 3 years. Up to 72% discount. Standard RI cannot change type. Convertible RI can change type.
Payment options: No Upfront (lowest discount), Partial Upfront, All Upfront (highest discount)Savings Plans: Commit to a spend in dollars per hour. More flexible than Reserved Instances.
Compute Savings Plans → any EC2 instance type, region, Fargate, LambdaEC2 Instance Savings Plans → specific instance family and region, more discountSpot Instances: Up to 90% off On-Demand. AWS can reclaim with 2-minute warning.
Spot request types: One-time → launches once, if terminated request closes Persistent → if terminated, request reopens and launches a new instanceRememberTo stop a persistent Spot setup completely, cancel the Spot Request FIRST then terminate the instance. Terminating alone causes the persistent request to reopen and launch a new instance.
Dedicated Hosts: Entire physical server for your use only. Required for BYOL (Bring Your Own License) compliance. Most expensive option.
Capacity Reservations: Guarantee capacity in a specific AZ. No pricing discount. Combine with Reserved Instances or Savings Plans to get discounts.
Decision guide: Short-term or unpredictable? → On-Demand Steady, running 24/7 for 1+ years? → Reserved or Savings Plans Fault-tolerant, can handle interruption? → Spot (up to 90% savings) BYOL licensing or compliance isolation? → Dedicated Hosts Guaranteed capacity in specific AZ? → Capacity ReservationsPlacement Groups
By default AWS decides where to physically place your EC2 instances. Placement Groups let you control that decision.
Cluster — all instances on same hardware in one AZ:
All packed together on same rackUltra-low latency and up to 10 Gbps between instancesIf the rack fails: all instances fail togetherUse for: HPC, ML training, big data jobs needing maximum network speedSpread — each instance on completely separate hardware:
Maximum fault isolation — each instance on different rack, different powerMaximum 7 instances per Availability ZoneUse for: critical instances where simultaneous failure is unacceptablePartition — instances divided into groups (partitions) on separate racks:
Each partition on separate rack hardwareScales to hundreds of instancesUse for: distributed systems — Hadoop, Cassandra, KafkaUp to 7 partitions per AZ| I need | Use |
|---|---|
| Maximum speed between instances | Cluster |
| Maximum fault isolation, few critical instances | Spread |
| Large-scale distributed system with some isolation | Partition |
Elastic Network Interfaces
An ENI (Elastic Network Interface) is a virtual network card. It contains the instance's IP address, MAC address, and Security Group associations. The key insight: these belong to the ENI, not the instance.
The failover use case:
Without ENI failover: App connects to database at: 10.0.1.50 Database instance crashes New instance launches → gets new IP: 10.0.1.99 App still pointing to 10.0.1.50 → connection fails Someone must update app config manually With ENI failover: ENI holds IP: 10.0.1.50, attached to Database Instance A Instance A crashes Detach ENI from Instance A Attach ENI to Instance B Instance B now has IP: 10.0.1.50 App connects to 10.0.1.50 → still works Zero config changes neededKey facts:
A custom ENI persists even after the instance it was attached to is terminatedMoving an ENI transfers the IP, MAC address, and Security Groups to the new instanceENIs are bound to a specific AZ — cannot move to a different AZEC2 Hibernate
Hibernate saves the contents of RAM to the root EBS volume. When the instance starts again, that RAM state is reloaded and the instance resumes exactly where it left off.
Normal stop: RAM data LOST, full OS restart on next boot (slow)Hibernate: RAM data SAVED to EBS, instance resumes from saved state (fast)Requirements:
Root volume must be EBS (not Instance Store)Root EBS volume MUST be encrypted (RAM may contain sensitive keys and tokens)RAM size must be less than 150 GBMaximum hibernate duration: 60 daysUse when:
Application has a long startup or initialisation timeLong-running calculation that cannot be interruptedResume the exact state at next business day without re-initialisationHands-on Lab — Launch EC2 with User Data and IAM Role
Step 1 — Create a Security Group
EC2 → Security Groups → Create security groupName: devops-web-sg VPC: your VPC Add inbound rules: SSH port 22 → My IP (your IP only — never 0.0.0.0/0) HTTP port 80 → Anywhere (0.0.0.0/0)Create security groupStep 2 — Create an IAM Role for EC2
IAM → Roles → Create roleTrusted entity: AWS service → EC2 → NextPermissions: AmazonS3ReadOnlyAccess → NextRole name: devops-ec2-s3-role → Create roleStep 3 — Launch EC2 with User Data
EC2 → Launch InstancesName: devops-web-01AMI: Amazon Linux 2023 (Free tier eligible)Instance type: t2.microKey pair: select your existing key pairSecurity groups: devops-web-sg Expand Advanced details:IAM instance profile → devops-ec2-s3-roleUser data → paste:yum update -yyum install -y nginxecho "<h1>DevOps Network — $(hostname -f)</h1>" > /usr/share/nginx/html/index.htmlsystemctl start nginxsystemctl enable nginxLaunch instanceStep 4 — Test the web server
Wait 2 minutes for User Data to finish runningCopy Public IPv4 address from instance detailsOpen browser → http://YOUR-PUBLIC-IPYou should see: DevOps Network — ip-10-x-x-x.ec2.internalStep 5 — Connect and test the IAM Role
EC2 → Instances → devops-web-01 → Connect → EC2 Instance Connect → Connect In the browser terminal:## Test S3 access — no aws configure needed, role provides credentialsaws s3 ls --region ap-south-1 ## See which identity is being usedaws sts get-caller-identity## Shows the role ARN — confirms role is active, no personal keys anywhereStep 6 — Test Security Group behaviour
EC2 → Security Groups → devops-web-sg → Inbound rules → EditDelete the HTTP port 80 rule → Save rules Try: http://YOUR-PUBLIC-IP → connection timeout (packet blocked) Add the rule back → Save rulesTry again → page loads immediately Security Group changes take effect instantly — no reboot needed.Step 7 — Observe IP change on stop/start
EC2 → Instances → devops-web-01 → Instance state → Stop instanceWait until stopped → Instance state → Start instanceCheck Public IPv4 address — it has changed to a new IPCheck Private IPv4 address — it is still the same This is why Elastic IP exists: to keep a stable public IP across stop/start.Step 8 — Cleanup
EC2 → Instances → devops-web-01 → Instance state → Terminate instanceEC2 → Security Groups → devops-web-sg → Actions → Delete security groupIAM → Roles → devops-ec2-s3-role → DeleteProduction Best Practices and Common Pitfalls
- Always use IAM Roles — never run aws configure on an EC2 instance and store access keys
- Use SSM Session Manager instead of SSH where possible — no port 22 needed, full audit trail
- Restrict SSH Security Group source to your specific IP — never 0.0.0.0/0 on port 22
- Pre-bake applications into custom AMIs — reduces boot time from minutes to seconds, critical for Auto Scaling
- Tag every EC2 instance from day one — Name, Environment, Team, Project — Cost Explorer is useless without tags
- Use Spot Instances for batch jobs, CI/CD build agents, and ML training — up to 90% savings
- Enable detailed monitoring on production instances — standard monitoring is only 5-minute intervals
Quick Reference and Troubleshooting Commands
| Task | Command |
|---|---|
| List instances | aws ec2 describe-instances --region ap-south-1 |
| Get instance public IP | aws ec2 describe-instances --instance-ids <id> --query 'Reservations[0].Instances[0].PublicIpAddress' --output text |
| Start instance | aws ec2 start-instances --instance-ids <id> --region ap-south-1 |
| Stop instance | aws ec2 stop-instances --instance-ids <id> --region ap-south-1 |
| Terminate instance | aws ec2 terminate-instances --instance-ids <id> --region ap-south-1 |
| Describe Security Groups | aws ec2 describe-security-groups --region ap-south-1 |
| List AMIs (yours) | aws ec2 describe-images --owners self --region ap-south-1 |
| List key pairs | aws ec2 describe-key-pairs --region ap-south-1 |
| Get instance metadata (from inside EC2) | curl http://169.254.169.254/latest/meta-data/instance-id |
Common problems and fixes:
| Problem | Likely cause | Fix |
|---|---|---|
| Connection timeout on SSH | Port 22 not open in Security Group | Add inbound rule for TCP 22 from your IP |
| Connection refused on SSH | Security Group is fine but SSH not running | Check if instance finished booting, verify SSH daemon running |
| Web server not reachable | Port 80 not open or nginx not started | Check Security Group inbound rule, verify nginx running via SSM |
| aws s3 ls returns no credentials | No IAM Role attached | Attach IAM Instance Profile with S3 permissions |
| User Data script did not run | Script has syntax errors or first boot already ran | Check /var/log/cloud-init-output.log for errors |
| Public IP changed after reboot | No Elastic IP assigned | Allocate and associate an Elastic IP if static IP is required |
Common MistakeOpening SSH (port 22) to 0.0.0.0/0 in the Security Group. Automated scanners hit every public IP on port 22 within minutes of an instance launching. Restrict SSH to your specific office or home IP. Use SSM Session Manager in production where port 22 never needs to be open at all.
TipWhen launching EC2 instances for Auto Scaling Groups, always create a custom AMI with your application pre-installed rather than relying on User Data to install everything at boot time. User Data installation takes 5-10 minutes per instance. A pre-baked AMI starts in 60-90 seconds. For a scale-out event adding 10 instances, the difference is 50-90 minutes of delayed capacity versus 10-15 minutes.
Common Mistakes to Avoid
Common MistakeOpening SSH (port 22) to 0.0.0.0/0 in the Security Group. Automated scanners try port 22 on every public IP in the world within minutes of launch. Restrict SSH to your specific IP only, or better, use SSM Session Manager with no port 22 at all.
Common MistakeStoring AWS access keys on an EC2 instance by running aws configure on the server. Use IAM Roles instead. The SDK fetches temporary credentials automatically. No keys stored anywhere. No risk.
TipPre-bake your application into a custom AMI instead of installing via User Data on every launch. Boot time drops from 8-10 minutes to 60-90 seconds. ASG reacts faster during scale-out events. This single change often cuts scaling overshoot by 60%.