Skip to main content

AWS Cost Optimisation: Cut Cloud Bills 40-60%

AWS bills surprise teams every month. Here are the 8 concrete actions that cut cloud spend by 40-60% without touching your application architecture.

A fintech startup in Mumbai was spending $28,000 per month on AWS. After a two-week audit, they discovered three things: development EC2 instances running 24/7 with 4% average CPU, S3 traffic routing through NAT Gateway instead of a free VPC Endpoint, and no Savings Plans despite running the same workload for 14 months. Three changes brought the bill to $16,000. No architecture changes. No application refactoring. Just fixing configuration decisions made in a hurry that nobody ever revisited.

AWS cost problems are almost always a collection of small waste points that accumulate over months. This guide walks through the eight highest-impact actions in order of effort and return.

Action 1: Rightsize EC2 Instances Before Buying Any Commitments

Buying Reserved Instances or Savings Plans on oversized instances locks in waste at a discount. Rightsize first, then commit.

Cost Explorer's Rightsizing Recommendations shows every EC2 instance running below 40% CPU utilisation over the past 14 days with a specific recommendation:

TEXT
Current: m5.2xlarge - $276/month On-Demand
Actual CPU: 8% average, 22% peak
Recommended: t3.large - $60/month On-Demand
Projected saving: $216/month on this one instance

A team with 20 oversized instances finding an average $80/month saving per instance saves $1,600/month, before buying any commitments.

Memory metrics require the CloudWatch Unified Agent on EC2. Enable it on every production instance:

Bash
## Install Unified CloudWatch Agent on Amazon Linux 2023
sudo yum install -y amazon-cloudwatch-agent
## Start with memory collection config
sudo /opt/aws/amazon-cloudwatch-agent/bin/amazon-cloudwatch-agent-ctl \
-a fetch-config -m ec2 \
-c ssm:/cloudwatch-config/prod \
-s

After rightsizing, your Savings Plan commitment is sized correctly — you are not locking in discount on compute you do not need.

Action 2: Buy Compute Savings Plans for Steady Workloads

On-Demand EC2 is the most expensive pricing option on AWS. Any workload that has been running for more than 2 months with a predictable pattern should have Savings Plans coverage.

Bash
## Check your current Savings Plans coverage
aws ce get-savings-plans-coverage \
--time-period Start=2024-01-01,End=2024-01-31 \
--granularity MONTHLY \
--region us-east-1

The recommendation engine in Cost Explorer does the work for you:

◈ DIAGRAM
Cost Explorer -> Savings Plans -> Recommendations
Term: 1-year Payment: No Upfront
Recommended commitment: $8.50/hour
Estimated annual savings: $42,000 (38% reduction)

Start conservative — commit to 80% of your baseline usage, not your peak. Spiky traffic above the commitment is charged at On-Demand rates, which is fine. The commitment covers your stable floor.

Compute Savings Plans apply automatically. When you switch instance types or regions, the discount follows. No manual reassignment like Reserved Instances require.

Action 3: Fix NAT Gateway Charges for S3 and DynamoDB

NAT Gateway charges $0.045 per GB processed. Every private EC2 instance accessing S3 through NAT Gateway pays this fee for traffic that never needed to touch the internet.

Bash
## Check what you are paying for VPC-related processing
aws ce get-cost-and-usage \
--time-period Start=2024-01-01,End=2024-01-31 \
--granularity MONTHLY \
--metrics BlendedCost \
--region us-east-1

The fix is a VPC Gateway Endpoint — free to create, instant, and reduces NAT costs for S3 and DynamoDB to zero:

Bash
## Create free S3 Gateway Endpoint
aws ec2 create-vpc-endpoint \
--vpc-id vpc-xxxxxxxx \
--service-name com.amazonaws.ap-south-1.s3 \
--route-table-ids rtb-private-1a rtb-private-1b \
--region ap-south-1
## Create free DynamoDB Gateway Endpoint
aws ec2 create-vpc-endpoint \
--vpc-id vpc-xxxxxxxx \
--service-name com.amazonaws.ap-south-1.dynamodb \
--route-table-ids rtb-private-1a rtb-private-1b \
--region ap-south-1

For a team with 15 TB/month of S3 traffic from private EC2:

TEXT
Without endpoint: 15,000 GB x $0.045 = $675/month in NAT charges
With endpoint: $0/month
Annual saving: $8,100

Two CLI commands. Eight thousand dollars per year saved.

Action 4: Tag Everything, Then Charge Teams for Their Usage

Without tags, Cost Explorer shows you what you are spending. With tags, it shows you who is spending it. The difference is the ability to hold teams accountable.

Mandatory tag policy:

Tag Key Example Value Purpose
Environment prod, staging, dev Split production from non-production costs
Team backend, frontend, data Charge costs to owning team
Project payments-v2, search-revamp Track project-level spend

Enforce tags with an AWS Config managed rule requiring the Environment and Team tags on every EC2 instance, so untagged resources get flagged automatically rather than discovered months later during a cost review.

Once tags are in place, use Cost Explorer to generate per-team cost reports monthly. When the backend team sees they spent $3,400 in dev last month, they become motivated to schedule dev instances correctly.

Action 5: Schedule Non-Production Instances Off Overnight

Development and staging instances are typically used during business hours only. Running them 24/7 wastes 65% of their cost.

Deploy AWS Instance Scheduler from the AWS Solutions Library as a single CloudFormation stack, set the default timezone to Asia/Kolkata, and tag each dev instance with a schedule:

Bash
## Tag dev instances to start at 8 AM, stop at 10 PM weekdays IST
aws ec2 create-tags \
--resources i-0abc123def456 \
--tags Key=Schedule,Value=india-business-hours \
--region ap-south-1
TEXT
Before scheduler: 20 dev EC2 instances x $0.10/hr x 720 hrs = $1,440/month
After scheduler (14 hrs x 22 working days): $0.10 x 308 hrs = $616/month
Saving: $824/month on dev instances alone

Action 6: Move S3 Data to the Right Storage Class

Teams frequently store everything in S3 Standard regardless of how often it is accessed. Logs written 90 days ago and never read again continue paying Standard prices.

Set lifecycle rules to move data automatically as it ages, transitioning to Standard-IA after 30 days and Glacier Flexible Retrieval after 90 days, with expiration after 365 days and cleanup of incomplete multipart uploads after 7 days. This is configured once per bucket through put-bucket-lifecycle-configuration and then runs without further intervention.

TEXT
S3 Standard: $0.023/GB/month
Standard-IA: $0.0125/GB/month (46% cheaper)
Glacier Flexible: $0.004/GB/month (83% cheaper)

A team storing 50 TB of logs that are rarely accessed after 30 days:

TEXT
All in Standard: 50,000 GB x $0.023 = $1,150/month
With lifecycle: mixed tiers based on age = ~$280/month
Saving: $870/month

Action 7: Use Spot Instances for Batch and Non-Critical Workloads

Spot Instances offer up to 90% savings on EC2. AWS can reclaim them with a 2-minute warning, making them unsuitable for stateful production services — but ideal for batch processing, data transformation, and CI/CD build agents.

Bash
## Launch Spot Fleet for nightly ETL, lowest-price allocation
aws ec2 request-spot-fleet \
--spot-fleet-request-config file://etl-spot-fleet.json \
--region ap-south-1

Specify multiple instance types in your Spot Fleet request — the more options you give AWS, the higher the probability of getting Spot capacity at any given time. For batch workloads that can checkpoint and resume, Spot Interruption Handling with a 2-minute warning handler is the standard pattern.

Action 8: Enable Cost Anomaly Detection

Cost anomaly detection uses ML to learn your normal spending patterns and alert you within hours when something unusual happens.

Create a dimensional anomaly monitor scoped to SERVICE covering all AWS services, then attach an email subscription with a $20 threshold and daily frequency so alerts land in a shared inbox rather than requiring anyone to check the console.

With a threshold of $20, you are alerted any time a service spends $20 more than expected in a day. A forgotten GPU instance, a misconfigured S3 replication rule, a runaway Lambda function — caught within 24 hours instead of at month-end invoice.

Trade-offs and Alternatives

Approach Effort Monthly Saving
Rightsize EC2 Medium $500-5,000
Savings Plans Low 30-72% of compute
NAT to VPC Endpoint Very low $200-2,000
Instance Scheduler Low $500-2,000
S3 Lifecycle rules Low $100-1,500
Spot for batch Medium 70-90% on batch
Cost Anomaly Detection Very low Prevents surprises

Production Implementation Guidelines

Execute in this order — each step builds on the previous:

  • Enable Cost Anomaly Detection immediately — this is free and catches future problems.
  • Tag all resources before anything else — you cannot optimise what you cannot attribute.
  • Rightsize EC2 using Cost Explorer recommendations before buying any commitments.
  • Buy Compute Savings Plans covering 80% of your steady baseline compute.
  • Create VPC Gateway Endpoints for S3 and DynamoDB in every VPC — free, instant.
  • Deploy Instance Scheduler for dev and staging environments.
  • Set S3 lifecycle rules on all log and archive buckets.
  • Move batch processing to Spot Instances starting with lowest-risk workloads.
Note

References and Further Reading

Frequently Asked Questions

Should you buy AWS Savings Plans before or after rightsizing EC2 instances?

Always rightsize first. Buying a Savings Plan commitment against oversized instances locks in a discount on wasted capacity — rightsize using Cost Explorer's recommendations, then size your Savings Plan commitment against the corrected baseline.

Why does NAT Gateway cost so much for S3 traffic, and how do you fix it without downtime?

NAT Gateway charges $0.045 per GB processed for any traffic routed through it, including S3 and DynamoDB calls that never needed to touch the internet. Creating a VPC Gateway Endpoint for S3 and DynamoDB is free, takes effect immediately, and requires no application changes or downtime.

What's the safest way to size a Compute Savings Plan commitment?

Commit to roughly 80% of your baseline steady-state usage, not your peak. Traffic above the commitment is simply billed at On-Demand rates, so under-committing costs less in flexibility than over-committing costs in wasted spend.

Are Spot Instances safe to use for production workloads?

Only for workloads that tolerate a 2-minute interruption warning and can checkpoint and resume — batch processing, ETL, and CI/CD build agents are good fits. Stateful production services that can't handle sudden reclaim are not.

How quickly can Cost Anomaly Detection catch a runaway resource like a forgotten GPU instance?

With a threshold around $20 and daily alert frequency, AWS's ML-based anomaly detection typically flags unusual per-service spend within 24 hours, compared to discovering it at month-end invoice review.

Discussion0