Cloud costs5 min read
Six AWS money leaks you can fix in an afternoon
Cloud bills rarely blow up because of one big decision. They blow up because of six small ones nobody revisits. Where to look, and what each costs.
When a company calls us because "the cloud got expensive", there is almost never one big decision behind it. There are six small ones, made two or three years ago, that nobody looked at again.
This is the list we check first in an audit. None of it requires rewriting code or migrating anything: these are configuration changes a team can make in an afternoon and see on the next invoice.
Prices are AWS list prices in us-east-1 as of September 2026. Swap in your own region before doing
the math: in São Paulo, for instance, almost everything costs more.
1. gp2 volumes that should be gp3
EBS volumes of type gp2 cost $0.10 per GB-month. gp3 volumes cost $0.08. Same disk, same
durability, 20% less.
The real gap is usually wider. With gp2, IOPS are tied to size: 3 IOPS per GB. To reach the IOPS
the database needed, many teams provisioned a 1 TB volume when 200 GB would have done. With gp3
you buy IOPS separately and get 3,000 as a baseline, so on top of the lower per-GB price you can
shrink the volume.
The change is live, with no instance restart. Start with volumes over 100 GB.
2. Public IPv4 addresses nobody uses
Since February 2024, AWS charges $0.005 per hour for every public IPv4 address, whether it is attached to anything or not. That is about $3.65 per IP per month.
It sounds like nothing until you count. Every instance with a public IP, every NAT Gateway, every load balancer, every Elastic IP left over from a project that was shut down in 2023. We have seen accounts with more than fifty live IPs where twelve were in use.
Public IP Insights, in the VPC console, gives you the list in two clicks.
3. A NAT Gateway processing traffic it shouldn't
A NAT Gateway costs $0.045 per hour plus $0.045 for every GB processed. The processed GB stacks on top of the internet egress charge, so traffic leaving a private subnet for the internet ends up costing around $0.135 per GB.
The problem is not the NAT — it is what goes through it. If your private instances talk to S3 or DynamoDB, that traffic has no business going out to the internet. VPC gateway endpoints for those two services are free and take that traffic off the NAT.
For workloads that move a lot of data to S3 (backups, ETL, images), this one change moves your network bill by an order of magnitude.
4. Orphaned volumes and snapshots
An EBS volume with no instance attached is billed exactly the same. So is a snapshot from four years ago.
The cause is almost always identical: someone terminated an instance without "delete on termination" checked, or a backup script that never had a retention policy. It is pure cost with nothing on the other side.
Before deleting anything: tag it, announce it, wait a week. An orphaned snapshot costs little; deleting the one somebody needed costs a lot.
5. Logs kept forever
By default, CloudWatch log groups have infinite retention. Nobody decides that; it ships that way.
The consequence is that you are paying storage for the debug logs of a service that was switched
off, and for every line your application wrote in debug mode during the month someone forgot to
turn it down.
Set retention per group: 7 days for debugging, 30 for application logs, whatever your compliance rules require for audit. If you need to keep more, export to S3 with lifecycle rules — storage there costs a fraction.
6. Everything at on-demand price
On-demand is the price of not committing. It is fine for anything you are not sure will still exist next month; it is expensive for anything that has been running non-stop for two years.
Two levers, easiest first:
- Savings Plans. You commit to an hourly spend for one or three years. Compute Savings Plans go up to 66% off and let you change family, size and region. EC2 Instance Savings Plans go up to 72%, but lock you to a family and a region.
- Graviton. AWS's ARM processors offer up to 40% better price-performance. If your workload is Python, Node, Java or Go and runs in containers, the move is usually rebuild and deploy.
| Leak | Where to look | Reference price |
|---|---|---|
| gp2 volumes | EC2 → Volumes, filter by type | $0.10 vs. $0.08 per GB-month |
| Public IPs | VPC → Public IP Insights | $0.005 per hour each |
| NAT Gateway | VPC → NAT Gateways, bytes metric | $0.045/hour + $0.045/GB |
| Orphans | EC2 → Volumes in "available" | Full volume price |
| Logs | CloudWatch → Log groups, retention column | Uncapped storage |
| On-demand | Cost Explorer → Recommendations | Up to 72% off |
How to measure it without arguing
Before you touch anything, take a snapshot. In Cost Explorer, group by usage type and look at the last three months. That table is what you will compare against afterwards, and it is what turns "I think it went down" into a number.
Two rules that save us trouble:
- Nothing gets switched off in week one. It gets tagged, measured and announced. Switching off fast is how production goes down, and how you lose the team's trust for round two.
- One change per deploy. If you move to gp3, set retention and buy Savings Plans on the same day, you will not know which of the three explained the difference.
None of these six leaks is an architecture mistake. They are adjustments nobody had time to revisit because the team was busy shipping features, which is their job. That is exactly why they work so well as a first step: they fix themselves in an afternoon and buy you the time for the big decisions.
- aws
- cost
- infrastructure