High load in 2026: the trends I’d actually design for — and the ones I’d ignore

Daniel Park

Daniel Park

September 18, 2026

High load in 2026: the trends I’d actually design for — and the ones I’d ignore

High load used to mean a number on a slide: requests per second, maybe concurrent connections if the audience was old enough to remember C10K. In 2026 the number that wakes me is rarely RPS. It is p99 while a deploy is rolling, it is the Postgres primary’s CPU while a “harmless” export runs, it is the bill for a queue that grew because a consumer held a transaction open across a Stripe call.

I still design for load. I ignore half the trends that get packaged as high-load architecture. This is the split I use when a team wants to “prepare for scale” and I want them to prepare for Tuesday.

What I mean by load now

Load is contention plus fan-out plus time. Contention is the lock, the sequence, the single hot partition, the one Redis key everyone increments. Fan-out is the one user action that becomes forty downstream calls because we grew microservices like a garden. Time is the tail: the request that waits on a cold JVM, a slow DNS, a GC, a tcp retransmit, a lock wait, then gets retried and does it twice.

If your architecture only talks about average RPS, you are designing a brochure. I want the shape of the worst minute: a deploy, a cache flush, a cron alignment, a client retry storm. That minute is the design.

I also want the cost shape. A system that holds 50k RPS on a fat Kubernetes cluster and a four-node Redis Enterprise bill is “high load” in a way that makes finance a stakeholder in every cache key. Cheap load is a feature. I have started putting a cost annotation on capacity docs. Engineers hate it. They hate a surprise invoice more.

Trends I will actually design for

Load shedding as a first-class path. When the primary is sad, I want a 503 with a Retry-After, a shed of the export job, a degrade of the recommendation call, not a thread pool that queues until the whole node is a latency bomb. Spring has been able to do this for years. Envoy and nginx can do it at the edge. Teams still skip it because it feels like giving up. Giving up on the non-critical path is how the critical path lives.

Timeouts that are budgets, not decorations. Every outbound call has a timeout that fits inside the caller’s budget. I draw the budget on the whiteboard: 200ms for the edge, 80ms for the pricing call, 40ms leftover for luck. If the numbers do not fit, I do not add a service. I add a cache or I drop a dependency from the request path. Virtual threads and goroutines make it easy to start more work. They do not create more milliseconds.

Queues with backpressure you can see. Kafka, SQS, NATS, a Postgres SKIP LOCKED table — I do not care, if the consumer can say “I am full” and the producer can slow down or drop with a metric. The 2018 playbook was “put it on the queue and scale consumers.” The 2026 failure is an infinite queue that hides a stuck consumer until the disk or the bill screams. I want lag as a page, and a poison-message story that is not “delete the topic.”

The database as a designed bottleneck. Postgres will be the wall. I design like it is. Connection pools sized to the primary, not to the number of pods. PgBouncer or equivalent in front. Indexes for the queries we run, not the ones we wrote in January. Read replicas for the questions that can be a second late. I do not design like a new Redis will save a bad query. Sometimes a cache helps. Sometimes it caches the bad plan’s result and makes the wrong thing fast.

Idempotency on anything that can retry. Clients retry. Envoys retry. Humans refresh. Stripe retries. If the handler is not idempotent, high load is how you double-charge. I want an idempotency key table or a unique constraint I can name in the spec. This is load design. It is also adult software.

Edge cache only where the contract allows it. Public GETs with ETags and a Vary I can explain. Private data stays private. I will not “add Cloudflare” to a authenticated API and hope. I have already written that autopsy. I will not write it again.

A monitoring wall showing latency heatmaps during a traffic spike

Trends I ignore on purpose

Just add Redis. Redis is a wonderful hammer. It is also a second availability domain, a second persistence story (or a dishonest one), and a second way to be single-threaded on the wrong key. I add it when I have a specific access pattern: a session, a rate limit, a lock with a TTL I understand, a cache I can lose. I do not add it because the last company had it. The playbook is dying because too many systems now have Redis as a mystery source of truth and a page that says OOM at 2am.

More microservices for isolation. Isolation is a real need. A process boundary is one way to get it. It is an expensive way if the only thing you isolated was a deploy calendar. I will split a service when the failure domain or the scaling axis is different — a CPU-heavy export versus a latency-sensitive checkout. I will not split a service so two teams can pretend they are decoupled while they share a database and a Slack argument.

Event-driven everything. Events are great for work that can be late. They are a way to lose a sale when used for work that cannot. I have seen “we will just emit OrderPlaced” become a scavenger hunt across six consumers, none of which could explain if the inventory reservation happened. If I cannot draw the success path as a sequence with a timeout, I do not call it an architecture. I call it a hope.

Multi-region as a default. Two regions is a product decision about who stays up when a zone dies, and a data decision about who is allowed to be wrong. It is not a badge. Active-active without a conflict story is how you get two inventories. I design single-region well first: multi-AZ, backups you have restored, a shed path. Then I ask which users pay for the second region.

AI at the request path. Inference on the checkout path is a latency and a reliability feature, not a demo. If the model is slow or sad, checkout must not be. I put model calls behind a budget and a default. “The recommendation is empty” is a valid degrade. “The card did not charge because the embedding endpoint 504’d” is not a 2026 trend I will honor.

Service meshes as a load strategy. A mesh can give you retries and mTLS. It can also multiply tail latency and hide timeouts in a sidecar you do not profile. I add a mesh when the security or the traffic policy needs it. I do not add it to “handle load.” Load is usually your query, your pool, or your retry storm.

A quiet on-call laptop showing a queue lag graph and a database connection pool

The first changes I make under real traffic

When a system is already hurting, I do not start with a rewrite. I start with a list I have used enough times to be boring:

  1. Find the top wait. pg_stat_activity, a mutex profile, a span that is fat. Not the dashboard tile that is green.
  2. Stop the retry storm. Cap client retries. Cap the mesh. Make the handler idempotent if it is not.
  3. Size the pool to the database, then size the pods to the pool. The other direction is how you get 4,000 connections and a primary that only does authentication.
  4. Move the export off the request path. The export is always on the request path. I do not know why. It is a law.
  5. Add a shed for the feature nobody will admit is optional. Usually search, recommendations, or a personalization call from a vendor.

Only after that will I talk about a new cache, a new topic, or a new region. Most teams want the new object because it feels like progress. The wait is progress.

A shape I would use for a checkout-ish service in 2026

One service for the request path. JVM or Go, I do not care as long as the timeouts are real. Postgres with a pooler. A queue for anything that talks to a payment provider after the authorize. Idempotency keys in a table with a unique constraint. Redis only if the rate limiter or the lock is proven, and it can vanish. Public catalog on a cacheable GET. Private cart uncached. OpenTelemetry traces that include the query text on the slow ones, sampled.

No mesh until we have two languages and a compliance reason. No second region until we have a restore drill and a written story about cart consistency. No LLM on the path. A feature flag to turn off the pretty things.

That shape is not trendy. It survives the worst minute. High load in 2026 is the worst minute, plus the bill, plus the retry. Design for those. Ignore the rest until they show up as a wait you can name.

If a trend cannot name the wait it removes, it is a purchase. I am trying to buy fewer of those.

More articles for you