AWS Networking Things to Know as a Backend Engineer
As a backend engineer I can build an API, containerize it, and point it at Postgres. Where it actually runs on AWS, and how traffic gets to it, used to be a black box. Going through Adrian Cantrill's SAA course fixed most of that.
This is not a full breakdown. These are the networking pieces I wish I had known before deploying anything.
A VPC is just your private network
Every resource you launch lives inside a VPC, a private network you define
with a CIDR range like 10.16.0.0/16.
Pick that range carefully. If you ever peer two VPCs or connect to an office
network, the ranges can't overlap, and changing it later means rebuilding. Avoid
10.0.0.0/16 since that is what everyone (and every default) uses.
Subnets live in one AZ
A VPC spans a whole region. A subnet lives in exactly one availability zone. So if you want your app to survive an AZ going down, you need subnets in at least two AZs and something spreading traffic across them.
Also, AWS reserves 5 IPs in every subnet. A /28 gives you 16 addresses on
paper and 11 you can actually use.
Public vs private is just a route
There is no "public" checkbox on a subnet. A subnet is public because its route
table sends 0.0.0.0/0 to an Internet Gateway. That's it.
| Subnet | Route for 0.0.0.0/0 | Who can reach it |
|---|---|---|
| Public | Internet Gateway | The internet (if the SG allows it) |
| Private | NAT Gateway | Nobody from outside, but it can call out |
| Isolated | none | Only things inside the VPC |
Fun detail: your EC2 instance never knows its own public IP. The OS only sees the private one, the Internet Gateway translates between them.
Where your stuff should go
The usual setup for an API:
- Load balancer: public subnets
- App servers / containers: private subnets
- Database (RDS): private or isolated subnets
Only the load balancer faces the internet. Your app and database should never have a public IP. Coming from Heroku or Railway, this is the biggest mental shift, the platform used to decide this for you.
NAT Gateways cost real money
Private subnets still need outbound internet to pull packages, call Stripe, or hit an external API. That is what a NAT Gateway is for. It sits in a public subnet and lets private resources call out without anyone calling in.
It charges per hour and per GB processed. For high availability you need one per AZ, so the bill multiplies.
Use a gateway endpoint for S3
A VPC endpoint lets resources reach AWS services without going over the internet. For S3 and DynamoDB there are gateway endpoints, and they are free. Add one and your S3 traffic skips the NAT Gateway entirely.
Other services (ECR, Secrets Manager, SQS) use interface endpoints, which cost per hour per AZ. Worth it at scale, maybe not for a side project.
Security groups vs NACLs
Both are firewalls, but they work differently.
| Security Group | NACL | |
|---|---|---|
| Attached to | A resource (ENI) | A subnet |
| State | Stateful | Stateless |
| Rules | Allow only | Allow and deny |
Stateful means if a request is allowed in, the response is allowed back out
automatically. NACLs are stateless, so you also have to allow the return
traffic on ephemeral ports (1024-65535), which is where people get stuck.
In practice you mostly live in security groups and leave NACLs on the default.
The best part: a security group can reference another security group instead of an IP.
| Security group | Inbound rule |
|---|---|
alb-sg | 443 from 0.0.0.0/0 |
app-sg | 8000 from alb-sg |
db-sg | 5432 from app-sg |
Now only your app can talk to Postgres, no matter how many instances scale up or what IPs they get.
Cross-AZ traffic isn't free
Data moving between AZs is billed in both directions. Usually that's fine, it is the price of high availability. Just know that a chatty app in one AZ talking to a database in another adds up.
Timeout vs connection refused
This one saves a lot of debugging time.
- Connection timed out: your packets never got an answer. Something in the network is dropping them, a security group, NACL, or route table.
- Connection refused: you reached the machine, but nothing is listening on that port. That's an app problem, not AWS.
And the classic refused case: the app is bound to localhost.
uvicorn app.main:app --host 127.0.0.1 --port 8000 # only reachable from inside the box
uvicorn app.main:app --host 0.0.0.0 --port 8000 # reachable from the load balancerThat's most of what matters day to day. VPCs, subnets, routes, and security groups cover nearly every "why can't my service reach X" problem I've hit. The rest of the networking section of the course is worth it, but these are the parts I use the most.