System design tips that actually matter in production
The habits I actually use on backends: clarify the problem, put numbers on it, name the trade-off, start simple, design for failure, and measure the bottleneck before adding Redis or Kafka.
System design tips that actually matter in production
I have sat in design reviews where the first drawing looked like this:
Client → API Gateway → Microservices → Kafka → Redis → Elasticsearch → Database
It looks expensive. It looks senior. Then I ask the only question that matters: why do we need all of this?
Most of the time the honest answer is “this is what a serious system looks like.” I have shipped NestJS APIs, Postgres, Redis, jobs, and production debugging. The systems that stayed calm were not the ones with the most boxes. They were the ones where every box had a constraint behind it.
Memorizing patterns is easy. The work is trade-offs under real traffic.
1. Clarify before you design
Jumping into architecture feels productive. It is how I have watched teams build the wrong system with a lot of confidence.
I start with verbs, not services. Create order. Capture payment. Show inventory. Refund. If a box does not serve a verb, it is decoration.
Then the constraints that actually change the drawing: traffic shape, latency, availability, durability, consistency. A profile I can serve 30 seconds stale is a cache. A balance that cannot be stale is not. Same product. Different API.
Read-heavy vs write-heavy is the other split I refuse to skip. A catalog that is 95% GET /product/:id on a small set of SKUs is a cache-shaped problem. An ingest path that must take a burst of events and query them later is a queue-shaped problem. Both “use a database.” They are not the same backend.
Before I draw, I want answers to a few questions:
What is the one flow that cannot be wrong — money, identity, inventory?
Is the hot path mostly reads or writes?
If two users update the same row, who wins?
What can wait a few seconds, and what must return in this request?
What happens when a dependency I do not own is slow or down?
If I cannot answer those, I am not ready to argue about sharding.
2. Capacity estimation: put numbers behind the design
I do not need a perfect model. I need the order of magnitude — “one Postgres is fine” vs “this will melt a single node.” That gap is huge. Most products I work on live in the first bucket longer than people want to admit.
10 million requests a day is 10,000,000 / 86,400 ≈ 116 QPS on average. Peak is often 5–10× that: 600–1,200 QPS. A well-indexed primary will often laugh at that if queries are cheap.
Attach size. 2 KB responses at 1,200 QPS is a trickle. 200 KB blobs on every request is a different conversation. 50 million rows at 500 bytes is still one database for a lot of teams. 5 million new rows a day is a retention problem, not a day-one Cassandra problem.
Those numbers tell me whether I need a cache, a replica, a queue, or just an index. “We might shard someday” is a growth path. It is not a cluster I have to operate this quarter.
Peak is the number that bites. I size for 9am, not the 24-hour mean. I also do not size for an imaginary 100× launch unless I have a reason.
The estimate can be wrong. It cannot be absent.
3. Trade-offs: say the cost out loud
This is the part that separates a design from a mood board. Almost every decision I like has a bill. For anything that matters I walk this:
Benefit → Downside → Mitigation → Alternative
I used to hear “put it in Redis, it will be fast.” Sometimes that is right. If the same product row is hit constantly, a cache drops database load. I now have two copies of the data. Invalidation is the design. TTL is blunt. Event-based delete is sharper and easier to get wrong. A cache stampede is what happens when a hot key expires and every Node process hits Postgres together. I have watched that look like a “database incident” when the database was fine.
Same for the rest:
Consistency vs availability: refuse the purchase if stock cannot be confirmed, or take the order and reconcile — oversell is a product decision, not a slogan.
Latency vs cost: extra replicas and regions make p99 prettier. An internal admin tool for 40 people does not need a global edge.
Complexity vs speed: microservices can let two teams deploy independently. They also mean timeouts, contracts, and an on-call graph. A modular NestJS app with clear modules is slower to tweet about and faster to debug at 2am.
Sync vs async: if “paid” must show before the page continues, capture is synchronous. If I also send the receipt in that same request, I have coupled money to SMTP. A queue helps — until consumers lag and nobody owns the alert.
SQL vs NoSQL: Postgres is my default because of transactions, joins, and EXPLAIN. A document store can fit a key-value shape I never join. It does not “scale better” as a personality.
A queue can buffer a spike. If producers outrun consumers all afternoon, lag and disk still grow. I stored the problem. I did not delete it.
If I cannot say the downside in a sentence, I have a favorite, not a design.
4. Start simple, then scale
I have seen engineers copy a FAANG diagram into a product that has 200 QPS and three people on-call. That is not ambition. That is unpaid operational work.
This is enough for a surprising number of APIs I have shipped:
Client → API → Database
I add a load balancer when there are two API instances. A cache when I have measured repeated reads. A queue when email or “notify warehouse” is taking checkout down. A replica when a report is stealing I/O from the money path. Partitioning when a table is large *and* queries have a natural key. Sharding when the primary cannot take writes after indexes and query shape are honest. Another service when a team or a scaling axis forces it — not when the diagram feels empty.
A path I have actually walked: ship the orders API on Postgres. Traffic grows, GET /orders/:id is most of QPS, add a short TTL and invalidate on update. Finance wants a scan of the whole table — send that to a replica. Black Friday writes spike because side effects sit in the request — move them to workers. Still one app. Still one primary for checkout.
Do not introduce Kafka because another company published an architecture post. Introduce it because a requirement, a bottleneck, or a failure mode showed up in *this* system.
5. Design for failure
I assume dependencies will fail, because they have: timeouts, a full disk, a payment provider in a bad region, Redis restarting empty, a queue that is up while the consumer is not.
Reliability is not preventing every failure. It is containing it so the important verbs still work.
Timeouts. Waiting forever is how one slow client drains the pool and the next healthy request waits behind the dead ones. Every outbound call needs a timeout I chose. “The default” is not a choice.
Retries. Useful for a blip. Dangerous if the call charges a card. Retry only when it is likely transient and the operation is idempotent — or keyed so the server can ignore a duplicate.
Backoff. Immediate retries from every Node process turn a blip into a stampede. Jitter is politeness and self-defense.
Circuit breakers. After consecutive failures, stop calling, fail fast, probe later. Useful when the downstream is clearly dead. Easy to get too eager.
Graceful degradation. If recommendations are on fire, checkout should still take payment. Dull is better than down. Rank the verbs. Protect the ones that are the product.
6. Observability is part of the design
If I cannot tell what the system is doing in production, I do not know how it behaves. I have a diagram and some hope.
Logs: what happened to this order id?
Metrics: is p95 up for everyone, or is this one tenant?
Traces: where did this request spend 800ms?
A correlation id on every hop is how those three talk to each other. Without it, “checkout is slow” is a scavenger hunt.
A pattern I keep seeing: someone reports checkout took ten seconds. I do not start by adding Redis. Metrics first — error rate on the payments client, pool saturation. Trace next — nine seconds in POST /payments/capture, 40ms in my database. Logs last — timeout, retry, then 200. That is a dependency, not a “backend is slow” story.
I want ids and red metrics on the money path before the first incident. After the incident I will add them in a hurry and miss the one request that mattered.
7. Optimize the actual bottleneck
“Let’s add Redis” is not a diagnosis. Neither is “let’s shard” or “let’s put Kafka in front.” Those are solutions looking for a graph I have not opened.
Measure → identify the bottleneck → understand the cause → change one thing → measure again.
I have watched “the database is slow” turn out to be a missing index, an N+1 to a user service, a 2 MB JSON blob, a pool of 10 connections, or a single counter row every checkout updates. I have watched CPU sit high on serialization, not on SQL. I have watched a third-party API in another region own p99 while we argued about cache TTLs.
Premature optimization, in my work, is not “writing tight code.” It is introducing a second datastore and a new failure mode because an 80ms p99 offended us on a 40 QPS admin API.
Complexity is easy to add and expensive to remove. I try to make the bottleneck introduce itself.
How I actually think through a design
Requirements → Numbers → Constraints → Trade-offs → Failure → Observability → Bottleneck
What must work. Rough QPS and size. What I cannot wish away. Benefit and cost out loud. What stays up when a friend dies. How I will know, with an id I can grep. Then the extra box.
Skip to the last step and I will install a queue in front of a missing index.
Interviews reward talking this through in 45 minutes. Production also charges for the bill, the deploy graph, and who owns the worker at 2am. A design the team cannot operate is not strong. It is expensive.
I do not try to know the most patterns. I try to know what problem I am solving, what the numbers imply, what can fail, and which trade-offs I am willing to live with.
The boxes are easy. The constraint is the job.