How Much Traffic Can a Single VPS Actually Handle?

Andre Okonkwo

Andre Okonkwo

September 27, 2026

How Much Traffic Can a Single VPS Actually Handle?

A few years ago a customer told me they were about to route their whole order flow through my API. Up to that point, my busiest day had been around 40,000 requests. Their estimate for launch week was “a few million a day, maybe more at peaks.” Everything ran on one VPS: 2 vCPUs, 4 GB of RAM, about $12 a month. My first instinct was to start pricing load balancers.

Instead I spent a weekend measuring what that box could actually do. The answer surprised me in both directions. The server could handle far more than I’d assumed on most endpoints, and one endpoint would have fallen over at about 16 requests per second no matter how big a machine I bought. The server size was never the real limit. How each request spent its time was.

So the honest answer to “how much traffic can a single VPS handle” is: it depends on what one request costs, and you can find out in an afternoon. Here’s how I did it and what I found.

Why “visitors per month” is the wrong unit

Hosting marketing talks in visitors or page views per month. Servers don’t experience months. They experience concurrent requests and requests per second at the peak.

Some arithmetic helps. A million requests a day averages out to about 12 per second. That sounds tiny, and it is. But traffic isn’t flat. Most sites see a daily peak several times the average, and a launch, a newsletter, or a link from a big site can produce a spike of ten times that for an hour. So a million requests a day might mean 12 per second at 4 a.m. and 100 or more per second at the worst moment.

Also, a page view is not one request. A page on a typical site loads HTML, CSS, JavaScript, fonts, and images, often 20 to 50 requests. If those static files are served by your VPS, they count. If a CDN serves them, only the HTML hits your server. For an API, one call is usually one request, which makes it easier to reason about.

What one request costs is the whole question

Here are rough numbers from that 2 vCPU box, measured with a load generator running on a separate machine in the same region. Your numbers will differ with language, framework, and hardware. The spread between them is the point.

  • A small static file served directly by the reverse proxy: many thousands of requests per second. The CPU barely noticed.
  • A JSON response from the app with no database call (a health endpoint, a cached config blob): around 2,000 to 3,000 per second before latency started climbing.
  • A typical API endpoint doing an auth check and two or three indexed database queries: roughly 400 to 600 per second.
  • A reporting endpoint that aggregated a month of data: about 20 per second.
  • The order-intake endpoint, which called a third-party address validation API before writing: about 16 per second, with the CPU nearly idle.

That’s a range of a thousand times between the fastest and slowest endpoints on the same server. The size of the machine barely matters by comparison.

Cars bunching up where a single-lane on-ramp merges into rush-hour traffic

The bottleneck that wasn’t CPU

That last number is the one that mattered, because it was exactly the endpoint the new customer would hit.

My app ran under a process server with a fixed number of synchronous workers: five, the usual formula for two cores. Each worker handles one request at a time. The address validation call took about 300 milliseconds. So each worker could complete about three of those requests per second, and five workers could do about 16. The sixth concurrent request waited in a queue. By the fiftieth, requests were timing out.

The CPU was at 10 percent during all of this. The workers weren’t computing anything. They were waiting on a network call. A bigger VPS with the same worker count would have done exactly 16 per second too.

This is the most common way small servers “run out of capacity” in my experience: not CPU, not RAM, but a fixed pool of workers or connections held open by something slow. The fixes are about the request, not the box:

  • Move slow external calls out of the request. I changed order intake to validate the payload, write it to the database with a “pending” status, return 202 Accepted, and let a background worker call the address API. The endpoint went from 16 per second to several hundred.
  • Use async workers for I/O-heavy endpoints if your stack supports it, so one process can wait on many network calls at once.
  • Set timeouts on every outbound call. A third-party API that hangs for 30 seconds will tie up workers much faster than one that fails quickly.

The next ceiling: database connections

After fixing order intake, my next instinct was to raise the worker count so the regular endpoints could take more load. Doubling workers roughly doubled throughput on the database-heavy endpoints, up to a point, and then errors appeared: the app couldn’t get a database connection. Each worker kept its own small pool, the background workers had theirs, and together they’d hit the database’s connection limit with the CPU still at 60 percent. At that point the question stops being about the VPS and becomes whether Postgres needs connection pooling in front of it. For me it didn’t yet; smaller per-worker pools were enough.

RAM, disk, and bandwidth

The other resources matter, but they usually show up later for a small app:

  • RAM limits how many workers you can run, and how much of your database fits in memory. When the working set of your data no longer fits in RAM, database queries start reading from disk and slow down sharply. On 4 GB with a database of a few hundred megabytes, I was nowhere near this.
  • Disk I/O matters for write-heavy workloads and for databases bigger than RAM. VPS disks vary a lot between providers and plans. A quick benchmark with fio tells you what you actually have.
  • Bandwidth is rarely the limit for APIs. My average response was about 2 KB; even 500 per second is only 1 MB per second. It matters more for sites serving images or downloads. Check both the port speed and the monthly transfer allowance, since some providers throttle or bill once you pass it.

Two laptops on a dim desk connected by an ethernet cable, one running hot

How to measure your own server in an afternoon

You don’t need to guess. This is the process I use now before any expected jump in traffic.

1. Test a copy, not production

Snapshot the server and spin up a copy of the same size. Load testing production will slow it down for real users, and some tests will break things. Point the copy at a copy of the database.

2. Generate load from a different machine

Running the load tool on the server you’re testing means they compete for the same CPU, and your numbers are meaningless. Use a second VPS in the same region. Tools like k6, wrk, oha, and hey all work; k6 is the easiest if you want to script a realistic mix of endpoints.

3. Test each important endpoint separately, then a realistic mix

Separate tests show you which endpoint breaks first. The mix shows you what the server does under a normal day’s traffic pattern. Weight the mix by what your logs say real users do.

4. Watch p95 latency, not averages

Increase the load in steps. At each step, record requests per second and the 95th percentile response time. For a while, throughput rises and latency stays flat. Then there’s a point where throughput stops rising and latency climbs steeply. That’s your real capacity. Averages hide this, because most requests are still fast while the slowest ones are timing out.

5. Watch the server while the test runs

Keep htop open, plus the database’s active connection count, plus your app’s worker or queue metrics. When latency climbs, look at what’s maxed out. If CPU is at 100 percent, you’re compute-bound and a bigger box helps. If CPU is low and latency is high, something is waiting: workers, connections, locks, or an external call. A bigger box won’t help that.

6. Leave headroom

Don’t plan to run at the capacity you measured. Queues behave badly as utilization approaches 100 percent: a server at 90 percent of capacity has much worse response times than one at 60 percent, and a small spike tips it over. I plan for normal peaks at about half of measured capacity.

Rough ballparks, with caveats

People always ask for a number, so here’s one, with the warning that your code matters more than any of it. For a small VPS with 2 to 4 vCPUs running a reasonably built app with a database on the same box:

  • A mostly static or well-cached site can serve millions of page views a day, especially with static assets on a CDN.
  • A typical CRUD app or API doing a few indexed queries per request can handle a few hundred requests per second at peak. That’s tens of millions of requests a day if the traffic is spread out.
  • Anything that does heavy computation per request, calls slow external services synchronously, or runs unindexed queries can fall over at tens of requests per second on any size of machine.

What happened with the customer

After those changes, the load test showed the box handling about 450 requests per second on a realistic mix with p95 under 150 ms. The customer’s launch week peaked at around 70 per second. CPU never went above 30 percent.

I did eventually move the database to its own server, about a year later, when its size outgrew the RAM. That was a data problem, not a traffic problem. The app still runs on one VPS. It’s a size bigger now, mostly because RAM got cheap, and the order-intake endpoint that would have failed at 16 requests per second is still the busiest thing on it.

If you’re worried about a traffic spike, measure before you buy anything. Most of the time the result is either “this box is fine” or “this one endpoint needs fixing.” Neither of those needs a second server.

More articles for you