All writing

AWS Networking Things to Know as a Backend Engineer

As a backend engineer I can build an API, containerize it, and point it at Postgres. Where it actually runs on AWS, and how traffic gets to it, used to be a black box. Going through Adrian Cantrill's SAA course fixed most of that.

This is not a full breakdown. These are the networking pieces I wish I had known before deploying anything.

A VPC is just your private network

Every resource you launch lives inside a VPC, a private network you define with a CIDR range like 10.16.0.0/16.

Pick that range carefully. If you ever peer two VPCs or connect to an office network, the ranges can't overlap, and changing it later means rebuilding. Avoid 10.0.0.0/16 since that is what everyone (and every default) uses.

Subnets live in one AZ

A VPC spans a whole region. A subnet lives in exactly one availability zone. So if you want your app to survive an AZ going down, you need subnets in at least two AZs and something spreading traffic across them.

Also, AWS reserves 5 IPs in every subnet. A /28 gives you 16 addresses on paper and 11 you can actually use.

Public vs private is just a route

There is no "public" checkbox on a subnet. A subnet is public because its route table sends 0.0.0.0/0 to an Internet Gateway. That's it.

SubnetRoute for 0.0.0.0/0Who can reach it
PublicInternet GatewayThe internet (if the SG allows it)
PrivateNAT GatewayNobody from outside, but it can call out
IsolatednoneOnly things inside the VPC

Fun detail: your EC2 instance never knows its own public IP. The OS only sees the private one, the Internet Gateway translates between them.

Where your stuff should go

The usual setup for an API:

  • Load balancer: public subnets
  • App servers / containers: private subnets
  • Database (RDS): private or isolated subnets

Only the load balancer faces the internet. Your app and database should never have a public IP. Coming from Heroku or Railway, this is the biggest mental shift, the platform used to decide this for you.

NAT Gateways cost real money

Private subnets still need outbound internet to pull packages, call Stripe, or hit an external API. That is what a NAT Gateway is for. It sits in a public subnet and lets private resources call out without anyone calling in.

It charges per hour and per GB processed. For high availability you need one per AZ, so the bill multiplies.

Use a gateway endpoint for S3

A VPC endpoint lets resources reach AWS services without going over the internet. For S3 and DynamoDB there are gateway endpoints, and they are free. Add one and your S3 traffic skips the NAT Gateway entirely.

Other services (ECR, Secrets Manager, SQS) use interface endpoints, which cost per hour per AZ. Worth it at scale, maybe not for a side project.

Security groups vs NACLs

Both are firewalls, but they work differently.

Security GroupNACL
Attached toA resource (ENI)A subnet
StateStatefulStateless
RulesAllow onlyAllow and deny

Stateful means if a request is allowed in, the response is allowed back out automatically. NACLs are stateless, so you also have to allow the return traffic on ephemeral ports (1024-65535), which is where people get stuck.

In practice you mostly live in security groups and leave NACLs on the default.

The best part: a security group can reference another security group instead of an IP.

Security groupInbound rule
alb-sg443 from 0.0.0.0/0
app-sg8000 from alb-sg
db-sg5432 from app-sg

Now only your app can talk to Postgres, no matter how many instances scale up or what IPs they get.

Cross-AZ traffic isn't free

Data moving between AZs is billed in both directions. Usually that's fine, it is the price of high availability. Just know that a chatty app in one AZ talking to a database in another adds up.

Timeout vs connection refused

This one saves a lot of debugging time.

  • Connection timed out: your packets never got an answer. Something in the network is dropping them, a security group, NACL, or route table.
  • Connection refused: you reached the machine, but nothing is listening on that port. That's an app problem, not AWS.

And the classic refused case: the app is bound to localhost.

uvicorn app.main:app --host 127.0.0.1 --port 8000  # only reachable from inside the box
uvicorn app.main:app --host 0.0.0.0 --port 8000    # reachable from the load balancer

That's most of what matters day to day. VPCs, subnets, routes, and security groups cover nearly every "why can't my service reach X" problem I've hit. The rest of the networking section of the course is worth it, but these are the parts I use the most.