Skip to main content

AWS VPC Security: Hardening Every Layer

Most cloud security incidents start with a misconfigured VPC. Here's how to harden every layer — subnets, Security Groups, NACLs, and IAM — for production.

A startup in Bengaluru left port 3306 open to 0.0.0.0/0 on their RDS Security Group. Not their EC2. Their RDS. The database was directly reachable from the internet. It took automated scanners 11 minutes to find it and begin brute-forcing. By the time the team noticed anomalous CloudTrail activity, 2.3 million user records had been exfiltrated.

VPC security is not a one-time setup. It is a layered architecture where each layer catches what the previous one misses. This guide walks through every layer — from subnet design to IAM conditions — the way production security teams at Razorpay and Zerodha implement it.

Why VPC Misconfiguration Is the Root Cause of Most AWS Breaches

The AWS Shared Responsibility Model puts network configuration entirely in your hands. AWS secures the physical infrastructure. You secure what runs on it.

The most common VPC mistakes that lead to breaches:

  • RDS, ElastiCache, or EC2 instances in public subnets with public IPs
  • Security Groups open to 0.0.0.0/0 on sensitive ports (22, 3306, 5432, 6379)
  • No VPC Flow Logs — no visibility when something goes wrong
  • Single NAT Gateway — availability risk and cross-AZ data transfer costs
  • No VPC Endpoints for S3 and DynamoDB — traffic going through the internet unnecessarily

None of these are difficult to fix. All of them are frequently found in production.

Subnet Architecture: Public and Private, Always Separate

The foundational rule: application servers and databases must never be in public subnets. The only resources in public subnets are load balancers, NAT Gateways, and Bastion Hosts.

TEXT
Correct production VPC layout:
VPC: 10.0.0.0/16
Public subnets (internet-facing):
10.0.1.0/24 - ap-south-1a - ALB, NAT Gateway
10.0.2.0/24 - ap-south-1b - ALB, NAT Gateway
Private subnets (application layer):
10.0.10.0/24 - ap-south-1a - EC2 app servers, ECS tasks
10.0.11.0/24 - ap-south-1b - EC2 app servers, ECS tasks
Private subnets (data layer):
10.0.20.0/24 - ap-south-1a - RDS primary, ElastiCache
10.0.21.0/24 - ap-south-1b - RDS standby, ElastiCache replica

Three tiers of private subnets give you defence in depth. Even if an attacker reaches the application layer, they cannot directly access the data layer — different Security Groups, different subnets, different route tables.

Disable auto-assign public IP on every private subnet. An EC2 instance accidentally launched in a private subnet with a public IP is a security incident waiting to happen.

Security Groups: The First Line of Defence

Security Groups are stateful allow-only firewalls attached to individual resources. They are the most frequently misconfigured element in AWS environments.

The Security Group chain pattern, enforce it everywhere:

TEXT
Internet
v allow 80, 443 from 0.0.0.0/0
ALB Security Group (sg-alb-prod)
v allow 8080 from sg-alb-prod only
App EC2 Security Group (sg-app-prod)
v allow 5432 from sg-app-prod only
RDS Security Group (sg-rds-prod)

Source is always a Security Group ID, never an IP address or a CIDR range. IPs change when instances restart. Security Group IDs never change. Using SG IDs as sources means your rules stay correct even as your fleet scales up and down.

The rule that must never exist in production:

Bash
## Port 22 open to the world - never do this
aws ec2 authorize-security-group-ingress \
--group-id sg-xxxxxxxx \
--protocol tcp --port 22 \
--cidr 0.0.0.0/0
## Correct: allow only from the app Security Group
aws ec2 authorize-security-group-ingress \
--group-id sg-rds-prod \
--protocol tcp --port 3306 \
--source-group sg-app-prod \
--region ap-south-1

NACLs: Subnet-Level Defence for Explicit Blocking

Network ACLs are stateless subnet-level firewalls. Unlike Security Groups, they support both allow and deny rules. They evaluate by rule number — lowest number wins.

The primary use case for NACLs: block known malicious IPs. If GuardDuty flags an IP as a threat, or your WAF identifies a DDoS source, add a NACL deny rule immediately.

Bash
## Block a known malicious IP at the subnet level
aws ec2 create-network-acl-entry \
--network-acl-id acl-xxxxxxxx \
--rule-number 10 \
--protocol -1 \
--rule-action deny \
--cidr-block 198.51.100.0/24 \
--ingress \
--region ap-south-1

The stateless trap: ephemeral ports.

NACLs must explicitly allow return traffic. When a client connects to your server on port 443, the response goes back on an ephemeral port (1024-65535) that the client chose. If your NACL blocks outbound traffic on the ephemeral range, the response never reaches the client.

Bash
## Allow HTTPS responses back to clients
aws ec2 create-network-acl-entry \
--network-acl-id acl-xxxxxxxx \
--rule-number 100 --protocol tcp \
--port-range From=443,To=443 \
--rule-action allow --cidr-block 0.0.0.0/0 \
--egress
## Allow ephemeral port range for return traffic
aws ec2 create-network-acl-entry \
--network-acl-id acl-xxxxxxxx \
--rule-number 110 --protocol tcp \
--port-range From=1024,To=65535 \
--rule-action allow --cidr-block 0.0.0.0/0 \
--egress

VPC Flow Logs: See Everything That Happens

Without VPC Flow Logs you are debugging network issues blind. Every accepted and rejected connection is recorded — source IP, destination IP, port, protocol, and whether it was accepted or rejected.

Bash
## Enable VPC Flow Logs to CloudWatch
aws ec2 create-flow-logs \
--resource-type VPC \
--resource-ids vpc-xxxxxxxx \
--traffic-type ALL \
--log-destination-type cloud-watch-logs \
--log-group-name /vpc/prod-flow-logs \
--deliver-logs-permission-arn arn:aws:iam::123456789012:role/FlowLogsRole \
--region ap-south-1

Debugging a mysterious connection failure:

◈ DIAGRAM
REJECT inbound -> Security Group or NACL blocking incoming traffic
REJECT outbound -> Security Group or NACL blocking outgoing traffic
ACCEPT in, REJECT out -> stateless NACL missing ephemeral port rule
ACCEPT both, app fails -> application-level issue, not network

At Razorpay, VPC Flow Logs feed into a CloudWatch Metrics Filter that alerts when there are more than 1,000 REJECT events per minute from a single source IP. This is the first indicator of a port scanning or brute-force attempt.

NAT Gateway: One Per AZ, Never One Total

A single NAT Gateway in one AZ creates two problems: it is a single point of failure for internet access, and instances in other AZs pay cross-AZ data transfer fees on every outbound request.

Bash
## Create NAT Gateway in ap-south-1a
aws ec2 allocate-address --domain vpc --region ap-south-1
aws ec2 create-nat-gateway \
--subnet-id subnet-public-1a \
--allocation-id eipalloc-xxxxxxxx \
--region ap-south-1
## Create a separate NAT Gateway in ap-south-1b
aws ec2 allocate-address --domain vpc --region ap-south-1
aws ec2 create-nat-gateway \
--subnet-id subnet-public-1b \
--allocation-id eipalloc-yyyyyyyy \
--region ap-south-1

Then associate each private subnet route table with the NAT Gateway in its own AZ. AZ-1a instances route through NAT GW A. AZ-1b instances route through NAT GW B. Neither depends on the other.

VPC Endpoints: Keep AWS Traffic Off the Internet

Private EC2 instances accessing S3 through a NAT Gateway pay NAT processing fees for traffic that never needed to touch the internet. A VPC Gateway Endpoint routes S3 and DynamoDB traffic directly through AWS's private network, free of charge.

Bash
## Create free S3 Gateway Endpoint
aws ec2 create-vpc-endpoint \
--vpc-id vpc-xxxxxxxx \
--service-name com.amazonaws.ap-south-1.s3 \
--route-table-ids rtb-private-1a rtb-private-1b \
--region ap-south-1
## Create free DynamoDB Gateway Endpoint
aws ec2 create-vpc-endpoint \
--vpc-id vpc-xxxxxxxx \
--service-name com.amazonaws.ap-south-1.dynamodb \
--route-table-ids rtb-private-1a rtb-private-1b \
--region ap-south-1

For a team reading 10 TB/month from S3 through a NAT Gateway, this single change saves roughly $450/month. Do it in every VPC from day one.

IAM as a Network Control: The Zero Trust Layer

Modern AWS security treats IAM as a network control. Even if traffic reaches an S3 bucket from inside your VPC, IAM decides whether the specific caller is allowed to access that specific object.

A bucket policy can deny all S3 access that does not originate from a specific VPC using a StringNotEquals condition on aws:SourceVpc. Even if someone has valid IAM credentials, they cannot access that bucket from outside the VPC — not from their laptop, not from another AWS account, not from any external system. See the AWS IAM policy reference (linked below) for the full condition-key syntax.

Similarly, a deny policy with BoolIfExists on aws:MultiFactorAuthPresent can require MFA for sensitive network-changing actions like modifying Security Group ingress rules or VPC attributes. Even if an attacker steals valid IAM credentials, they cannot open Security Group ports without the physical MFA device.

GuardDuty: Threat Detection Inside the VPC

GuardDuty monitors VPC Flow Logs, CloudTrail, and DNS logs automatically and uses ML to detect threats that rule-based systems miss.

Critical GuardDuty findings to wire up immediately:

Finding What It Means
UnauthorizedAccess:EC2/SSHBruteForce SSH brute force in progress
Recon:EC2/PortProbeUnprotectedPort Port scanning detected
CryptoCurrency:EC2/BitcoinTool.B Instance mining crypto
UnauthorizedAccess:IAMUser/TorIPCaller Credentials used via Tor

Wire findings to EventBridge, then to Lambda, for an automated response:

TEXT
GuardDuty finding: SSH brute force from a flagged IP
v
EventBridge rule matches
v
Lambda adds NACL deny rule for the IP automatically
Lambda sends a Slack alert to the security channel
Total response time: under 30 seconds

Production Implementation Guidelines

Implement these in order — each builds on the previous.

  • Deploy a three-tier VPC on day one: public (LB only), private-app, private-data. Never add this layer later — retrofitting subnets is painful.
  • Enforce Security Group chaining using SG IDs as sources. Audit weekly using AWS Config rules for restricted SSH and authorized-ports-only.
  • Enable VPC Flow Logs from the start. Storage cost is minimal. The debugging and incident response value is enormous.
  • Create S3 and DynamoDB Gateway Endpoints in every VPC immediately — free, faster, and eliminates unnecessary NAT traffic.
  • Enable GuardDuty in every region and every account. Wire findings to EventBridge for automated response.
  • Use AWS Config rules to continuously verify your security posture, not just at setup time.
Note

References and Further Reading

Frequently Asked Questions

Why is a Security Group ID safer as a source than an IP address or CIDR range?

IP addresses change whenever instances restart or scale, silently breaking rules that reference them. Security Group IDs stay fixed regardless of fleet changes, so chaining SG-to-SG references keeps your rules correct as instances scale up and down.

What causes "ACCEPT inbound, REJECT outbound" in VPC Flow Logs when the Security Group looks correctly configured?

This pattern almost always points to a stateless NACL missing the ephemeral port range (1024-65535) on egress — the response traffic to an inbound connection uses a client-chosen ephemeral port that the NACL is blocking, even though the Security Group correctly allowed the original request.

Do you need both Security Groups and NACLs, or is one enough?

Use both. Security Groups are stateful and resource-level; NACLs are stateless and subnet-level and are the only one of the two that supports explicit deny rules, making them the right tool for quickly blocking a known-malicious IP without touching every resource's Security Group.

Can IAM permissions substitute for VPC network isolation?

They complement it, not replace it. A bucket policy with an `aws:SourceVpc` condition can deny all access that doesn't originate from a specific VPC even with valid IAM credentials — treating IAM as an additional network-aware control layer on top of, not instead of, subnet and Security Group design.

What's the single most common root cause of AWS data breaches involving VPCs?

A database or cache resource left in a public subnet with a public IP and a Security Group open to 0.0.0.0/0 on a sensitive port — RDS, ElastiCache, and Redis instances reachable directly from the internet are consistently the pattern behind large-scale exfiltration incidents.

Discussion0