Skip to main content
Calcimator

Server Sizing Calculator

Calculate CPU, RAM, storage, and instance count from workload requirements.

About this calculator

This calculator sizes compute, memory, storage, and instance count from application-level traffic assumptions rather than raw hardware specs. It first derives requests per second from concurrent users and their per-minute request rate, then applies a form of Little's Law — average response time tells you how many requests are in flight at any instant (concurrency = throughput × latency). Each in-flight request is assumed to occupy about 0.15 vCPU, a moderate compute-intensity figure meant for typical web/API workloads rather than CPU-bound processing, and the raw CPU need is then inflated so the system runs at your target utilization ceiling (70% is the common production target, leaving headroom for spikes) rather than pegged at 100%. Memory sizing adds a fixed per-connection cost to the application's base memory footprint and layers on a 20% buffer.

Storage combines per-user data with a flat 20GB allowance for the OS and application itself. Instance count assumes a typical 8-vCPU instance size and always recommends at least two instances for high availability, even if the math would call for less. A "peak vCPUs" figure separately models what happens at double your normal traffic, so you can see the gap between steady-state and worst-case provisioning. Because the per-request CPU cost and per-instance vCPU count are fixed assumptions rather than measured from your actual workload, confirm these figures against a real load test before committing to a reserved-instance purchase.

Inputs

%

Results

Recommended vCPUs

1

Recommended RAM (GB)

2

Instances needed (HA)

2

Total storage (GB)520
Total requests/sec16.67
Concurrent requests3.3
Est. bandwidth (Mbps)6.51
Peak vCPUs (2x traffic)2
How to Use This Calculator
  1. Enter Concurrent Users, Requests per User per Minute, and Average Response Time (ms).
  2. Set Target CPU Utilization % to leave headroom for traffic spikes.
  3. Input Memory per Connection (MB) and App Base Memory (MB) to size RAM.
  4. Enter Storage per User (GB) and Total Registered Users.
  5. Review Recommended vCPUs, RAM (GB), and Storage (GB) to select the right server or cloud instance.

What each input means

Concurrent users
Maximum simultaneous active users on the system.
Requests per user per min
Average API/page requests each active user generates per minute.
Avg response time (ms)
Average server-side response time in milliseconds.
Target CPU utilization %
Target CPU utilization ceiling (70% is typical for production).
Memory per connection (MB)
Memory consumed per active user connection/session.
App base memory (MB)
Fixed memory used by the application (JVM heap, runtime, caches).
Storage per user (GB)
Average disk storage per registered user (uploads, data).
Total registered users
Total user accounts (for storage calculation).

What each result means

Recommended vCPUs
Total vCPUs needed across all instances.
Recommended RAM (GB)
Total RAM needed with 20% buffer.
Instances needed (HA)
Number of server instances for high availability (minimum 2).
Total storage (GB)
Disk storage for application data + OS.
Total requests/sec
Calculated requests per second at given load.
Concurrent requests
Average number of in-flight requests at any moment.
Est. bandwidth (Mbps)
Estimated network bandwidth assuming ~50KB avg response.
Peak vCPUs (2x traffic)
vCPUs needed to handle double the normal traffic.

How this is calculated

Worked example, using the default values

  1. Identify Input Parameters
    4 parameters
    Concurrent users = 500, Requests per user per min = 2, Avg response time (ms) = 200, Target CPU utilization % = 70 = 8 input(s) provided
  2. Calculate Recommended vCPUs
    Recommended vCPUs = max(1
    1 = 1
  3. Calculate Recommended RAM
    Recommended RAM = max(1
    2 = 2
  4. Calculate Instances needed
    Instances needed
    2 = 2
  5. Calculate Total storage
    Total storage
    520 = 520
  6. Calculate Total requests/sec
    Total requests/sec = (concurrentUsers * avgRequestsPerUserPerMin) / 60
    16.67 = 16.67

Engine last updated . Checked against 2 independently-derived tests — how we verify calculators. Built by Paul Gunder, a software engineer, not a licensed financial, medical, or legal professional.

Frequently Asked Questions

Why does lowering the target CPU utilization percentage increase the recommended vCPUs?

Target CPU Utilization % is the denominator the calculator divides your raw CPU need by, so a lower target — meaning you want more idle headroom — produces a larger recommendedVCPUs figure for the same traffic. Setting it to 50% instead of 70%, for example, provisions roughly 40% more CPU for identical concurrent-request load, trading cost for a bigger buffer against traffic spikes.

Why is the minimum instance count always 2, even for small workloads?

The instancesNeeded output is calculated from recommendedVCPUs divided by a typical 8-vCPU instance size, but the formula always returns at least 2 regardless of how small that division comes out — this models the standard high-availability practice of never running a production service on a single instance, since one node has no failover if it goes down.

What does the 'Peak vCPUs (2x traffic)' figure assume that the main recommendation doesn't?

The main recommendedVCPUs is sized for your entered concurrent users and target utilization as-is. Peak vCPUs reruns the same CPU-per-request math against double your calculated requests-per-second (peakRps = totalRps × 2) to show what capacity a traffic surge would demand — it's a separate 'what if' figure, not included in the main recommendation, so you'd need to provision toward it explicitly if you want standing headroom for a 2x spike.

How does the calculator turn concurrent users into a CPU requirement?

It works backward through a form of Little's Law: concurrent users and their requests-per-minute rate produce requests per second, which combined with your average response time gives the number of requests in flight at any instant (concurrency = throughput × latency). Each of those in-flight requests is assumed to cost about 0.15 vCPU, a fixed compute-intensity figure meant for typical web/API work rather than CPU-heavy processing like video encoding or ML inference.

The questions that sit next to this one — chosen by subject, including calculators filed under a different category.

More in Technology & Computing.