AWS VPC Security: Hardening Every Layer
Most cloud security incidents start with a misconfigured VPC. Here's how to harden every layer — subnets, Security Groups, NACLs, and IAM — for production.
A startup in Bengaluru left port 3306 open to 0.0.0.0/0 on their RDS Security Group. Not their EC2. Their RDS. The database was directly reachable from the internet. It took automated scanners 11 minutes to find it and begin brute-forcing. By the time the team noticed anomalous CloudTrail activity, 2.3 million user records had been exfiltrated.
VPC security is not a one-time setup. It is a layered architecture where each layer catches what the previous one misses. This guide walks through every layer — from subnet design to IAM conditions — the way production security teams at Razorpay and Zerodha implement it.
Why VPC Misconfiguration Is the Root Cause of Most AWS Breaches
The AWS Shared Responsibility Model puts network configuration entirely in your hands. AWS secures the physical infrastructure. You secure what runs on it.
The most common VPC mistakes that lead to breaches:
- RDS, ElastiCache, or EC2 instances in public subnets with public IPs
- Security Groups open to 0.0.0.0/0 on sensitive ports (22, 3306, 5432, 6379)
- No VPC Flow Logs — no visibility when something goes wrong
- Single NAT Gateway — availability risk and cross-AZ data transfer costs
- No VPC Endpoints for S3 and DynamoDB — traffic going through the internet unnecessarily
None of these are difficult to fix. All of them are frequently found in production.
Subnet Architecture: Public and Private, Always Separate
The foundational rule: application servers and databases must never be in public subnets. The only resources in public subnets are load balancers, NAT Gateways, and Bastion Hosts.
Correct production VPC layout:VPC: 10.0.0.0/16 Public subnets (internet-facing): 10.0.1.0/24 - ap-south-1a - ALB, NAT Gateway 10.0.2.0/24 - ap-south-1b - ALB, NAT Gateway Private subnets (application layer): 10.0.10.0/24 - ap-south-1a - EC2 app servers, ECS tasks 10.0.11.0/24 - ap-south-1b - EC2 app servers, ECS tasks Private subnets (data layer): 10.0.20.0/24 - ap-south-1a - RDS primary, ElastiCache 10.0.21.0/24 - ap-south-1b - RDS standby, ElastiCache replicaThree tiers of private subnets give you defence in depth. Even if an attacker reaches the application layer, they cannot directly access the data layer — different Security Groups, different subnets, different route tables.
Disable auto-assign public IP on every private subnet. An EC2 instance accidentally launched in a private subnet with a public IP is a security incident waiting to happen.
Security Groups: The First Line of Defence
Security Groups are stateful allow-only firewalls attached to individual resources. They are the most frequently misconfigured element in AWS environments.
The Security Group chain pattern, enforce it everywhere:
Internet v allow 80, 443 from 0.0.0.0/0ALB Security Group (sg-alb-prod) v allow 8080 from sg-alb-prod onlyApp EC2 Security Group (sg-app-prod) v allow 5432 from sg-app-prod onlyRDS Security Group (sg-rds-prod)Source is always a Security Group ID, never an IP address or a CIDR range. IPs change when instances restart. Security Group IDs never change. Using SG IDs as sources means your rules stay correct even as your fleet scales up and down.
The rule that must never exist in production:
## Port 22 open to the world - never do thisaws ec2 authorize-security-group-ingress \ --group-id sg-xxxxxxxx \ --protocol tcp --port 22 \ --cidr 0.0.0.0/0 ## Correct: allow only from the app Security Groupaws ec2 authorize-security-group-ingress \ --group-id sg-rds-prod \ --protocol tcp --port 3306 \ --source-group sg-app-prod \ --region ap-south-1NACLs: Subnet-Level Defence for Explicit Blocking
Network ACLs are stateless subnet-level firewalls. Unlike Security Groups, they support both allow and deny rules. They evaluate by rule number — lowest number wins.
The primary use case for NACLs: block known malicious IPs. If GuardDuty flags an IP as a threat, or your WAF identifies a DDoS source, add a NACL deny rule immediately.
## Block a known malicious IP at the subnet levelaws ec2 create-network-acl-entry \ --network-acl-id acl-xxxxxxxx \ --rule-number 10 \ --protocol -1 \ --rule-action deny \ --cidr-block 198.51.100.0/24 \ --ingress \ --region ap-south-1The stateless trap: ephemeral ports.
NACLs must explicitly allow return traffic. When a client connects to your server on port 443, the response goes back on an ephemeral port (1024-65535) that the client chose. If your NACL blocks outbound traffic on the ephemeral range, the response never reaches the client.
## Allow HTTPS responses back to clientsaws ec2 create-network-acl-entry \ --network-acl-id acl-xxxxxxxx \ --rule-number 100 --protocol tcp \ --port-range From=443,To=443 \ --rule-action allow --cidr-block 0.0.0.0/0 \ --egress ## Allow ephemeral port range for return trafficaws ec2 create-network-acl-entry \ --network-acl-id acl-xxxxxxxx \ --rule-number 110 --protocol tcp \ --port-range From=1024,To=65535 \ --rule-action allow --cidr-block 0.0.0.0/0 \ --egressVPC Flow Logs: See Everything That Happens
Without VPC Flow Logs you are debugging network issues blind. Every accepted and rejected connection is recorded — source IP, destination IP, port, protocol, and whether it was accepted or rejected.
## Enable VPC Flow Logs to CloudWatchaws ec2 create-flow-logs \ --resource-type VPC \ --resource-ids vpc-xxxxxxxx \ --traffic-type ALL \ --log-destination-type cloud-watch-logs \ --log-group-name /vpc/prod-flow-logs \ --deliver-logs-permission-arn arn:aws:iam::123456789012:role/FlowLogsRole \ --region ap-south-1Debugging a mysterious connection failure:
REJECT inbound -> Security Group or NACL blocking incoming trafficREJECT outbound -> Security Group or NACL blocking outgoing trafficACCEPT in, REJECT out -> stateless NACL missing ephemeral port ruleACCEPT both, app fails -> application-level issue, not networkAt Razorpay, VPC Flow Logs feed into a CloudWatch Metrics Filter that alerts when there are more than 1,000 REJECT events per minute from a single source IP. This is the first indicator of a port scanning or brute-force attempt.
NAT Gateway: One Per AZ, Never One Total
A single NAT Gateway in one AZ creates two problems: it is a single point of failure for internet access, and instances in other AZs pay cross-AZ data transfer fees on every outbound request.
## Create NAT Gateway in ap-south-1aaws ec2 allocate-address --domain vpc --region ap-south-1aws ec2 create-nat-gateway \ --subnet-id subnet-public-1a \ --allocation-id eipalloc-xxxxxxxx \ --region ap-south-1 ## Create a separate NAT Gateway in ap-south-1baws ec2 allocate-address --domain vpc --region ap-south-1aws ec2 create-nat-gateway \ --subnet-id subnet-public-1b \ --allocation-id eipalloc-yyyyyyyy \ --region ap-south-1Then associate each private subnet route table with the NAT Gateway in its own AZ. AZ-1a instances route through NAT GW A. AZ-1b instances route through NAT GW B. Neither depends on the other.
VPC Endpoints: Keep AWS Traffic Off the Internet
Private EC2 instances accessing S3 through a NAT Gateway pay NAT processing fees for traffic that never needed to touch the internet. A VPC Gateway Endpoint routes S3 and DynamoDB traffic directly through AWS's private network, free of charge.
## Create free S3 Gateway Endpointaws ec2 create-vpc-endpoint \ --vpc-id vpc-xxxxxxxx \ --service-name com.amazonaws.ap-south-1.s3 \ --route-table-ids rtb-private-1a rtb-private-1b \ --region ap-south-1 ## Create free DynamoDB Gateway Endpointaws ec2 create-vpc-endpoint \ --vpc-id vpc-xxxxxxxx \ --service-name com.amazonaws.ap-south-1.dynamodb \ --route-table-ids rtb-private-1a rtb-private-1b \ --region ap-south-1For a team reading 10 TB/month from S3 through a NAT Gateway, this single change saves roughly $450/month. Do it in every VPC from day one.
IAM as a Network Control: The Zero Trust Layer
Modern AWS security treats IAM as a network control. Even if traffic reaches an S3 bucket from inside your VPC, IAM decides whether the specific caller is allowed to access that specific object.
A bucket policy can deny all S3 access that does not originate from a specific VPC using a StringNotEquals condition on aws:SourceVpc. Even if someone has valid IAM credentials, they cannot access that bucket from outside the VPC — not from their laptop, not from another AWS account, not from any external system. See the AWS IAM policy reference (linked below) for the full condition-key syntax.
Similarly, a deny policy with BoolIfExists on aws:MultiFactorAuthPresent can require MFA for sensitive network-changing actions like modifying Security Group ingress rules or VPC attributes. Even if an attacker steals valid IAM credentials, they cannot open Security Group ports without the physical MFA device.
GuardDuty: Threat Detection Inside the VPC
GuardDuty monitors VPC Flow Logs, CloudTrail, and DNS logs automatically and uses ML to detect threats that rule-based systems miss.
Critical GuardDuty findings to wire up immediately:
| Finding | What It Means |
|---|---|
| UnauthorizedAccess:EC2/SSHBruteForce | SSH brute force in progress |
| Recon:EC2/PortProbeUnprotectedPort | Port scanning detected |
| CryptoCurrency:EC2/BitcoinTool.B | Instance mining crypto |
| UnauthorizedAccess:IAMUser/TorIPCaller | Credentials used via Tor |
Wire findings to EventBridge, then to Lambda, for an automated response:
GuardDuty finding: SSH brute force from a flagged IP vEventBridge rule matches vLambda adds NACL deny rule for the IP automaticallyLambda sends a Slack alert to the security channelTotal response time: under 30 secondsProduction Implementation Guidelines
Implement these in order — each builds on the previous.
- Deploy a three-tier VPC on day one: public (LB only), private-app, private-data. Never add this layer later — retrofitting subnets is painful.
- Enforce Security Group chaining using SG IDs as sources. Audit weekly using AWS Config rules for restricted SSH and authorized-ports-only.
- Enable VPC Flow Logs from the start. Storage cost is minimal. The debugging and incident response value is enormous.
- Create S3 and DynamoDB Gateway Endpoints in every VPC immediately — free, faster, and eliminates unnecessary NAT traffic.
- Enable GuardDuty in every region and every account. Wire findings to EventBridge for automated response.
- Use AWS Config rules to continuously verify your security posture, not just at setup time.
NoteReferences and Further Reading
- AWS VPC Security Best Practices — official AWS guidance
- AWS GuardDuty findings types — full catalogue of threats detected
- AWS Config managed rules — compliance rule library
Frequently Asked Questions
Why is a Security Group ID safer as a source than an IP address or CIDR range?
IP addresses change whenever instances restart or scale, silently breaking rules that reference them. Security Group IDs stay fixed regardless of fleet changes, so chaining SG-to-SG references keeps your rules correct as instances scale up and down.
What causes "ACCEPT inbound, REJECT outbound" in VPC Flow Logs when the Security Group looks correctly configured?
This pattern almost always points to a stateless NACL missing the ephemeral port range (1024-65535) on egress — the response traffic to an inbound connection uses a client-chosen ephemeral port that the NACL is blocking, even though the Security Group correctly allowed the original request.
Do you need both Security Groups and NACLs, or is one enough?
Use both. Security Groups are stateful and resource-level; NACLs are stateless and subnet-level and are the only one of the two that supports explicit deny rules, making them the right tool for quickly blocking a known-malicious IP without touching every resource's Security Group.
Can IAM permissions substitute for VPC network isolation?
They complement it, not replace it. A bucket policy with an `aws:SourceVpc` condition can deny all access that doesn't originate from a specific VPC even with valid IAM credentials — treating IAM as an additional network-aware control layer on top of, not instead of, subnet and Security Group design.
What's the single most common root cause of AWS data breaches involving VPCs?
A database or cache resource left in a public subnet with a public IP and a Security Group open to 0.0.0.0/0 on a sensitive port — RDS, ElastiCache, and Redis instances reachable directly from the internet are consistently the pattern behind large-scale exfiltration incidents.
Discussion0