Small rack of sage compute modules with one orange indicator: sizing the box count
AI Performance Engineering
QA / TESTING · AI PERFORMANCE ENGINEERING

Capacity planning and GPU sizing

How many boxes, which model, at what cost per unit of work. Measured, then extrapolated, never guessed.

THE PROBLEM

Where sizing numbers usually come from, and why they are wrong

Vendor throughput figures assume short prompts, warm caches, and ideal batching. Your workload has none of those.

The gap between the marketing number and the measured number is routinely large, and it is the difference between the right purchase and a second purchase six months later.

FIRST STEP

Pre-sizing on a sample

What it is: you provide a sanitized or synthetic sample of your workload (prompts, documents, or images) and the candidate models or serving stacks. We run it on our own known hardware and return tokens per second (or items per hour), the saturation point, and a GPU count range for your target volume, before you commit capital.

Scope: a paid first step with a fixed scope and a dated report; credited against a larger engagement when one follows.

Four Walls note: for clients whose data cannot leave, the same method runs on their hardware as the first phase of a load-test engagement.

DELIVERABLE: SIZING MEMO
  • [·]Throughput per device.
  • [·]Saturation point.
  • [·]GPU count range for the target volume.
  • [·]Cost per unit of throughput.
  • [·]The assumptions, stated.
FULL SIZING

Full sizing from load tests

When the system exists, sizing comes from the load tests described on the LLM and pipeline pages: measured curves at each concurrency, the breakpoint, and the resource that caused it, extrapolated to the target volume with a stated safety margin. Includes model and serving-engine comparisons at equal load, and quantization trade-offs measured for both throughput and output quality.

INPUTS

What goes into the number

SIZING INPUTS
  • [·]Prompt and output length distribution.
  • [·]Concurrency pattern (steady, bursty, batch windows).
  • [·]Cache behavior.
  • [·]SLO for TTFT and ITL, or for items per hour.
  • [·]Goodput target.
  • [·]Growth assumptions.
  • [·]The cost model (capital, cloud, or hybrid).
METHOD

Two layers, one report

Layer 1: application path (your load tool)

Scripts that drive the application the way users or upstream systems do, including streaming responses (server-sent events), with time to first token, per-token latency, and tokens per second captured per virtual user. We work in the load tool you already run: OpenText LoadRunner, Tricentis NeoLoad, Apache JMeter, Grafana k6, Gatling, or similar. If you have none, we bring the tooling.

Layer 2: inference path (AIPerf + engine metrics)

Direct benchmarking of the serving endpoint with NVIDIA AIPerf, correlated with the engine's own metrics (vLLM, llama.cpp, or the client's stack): TTFT, inter-token latency, tokens per second, request queueing, and KV-cache pressure.

Why both: the application path shows what users experience and where the app stack adds latency; the inference path shows where the model and hardware actually run out. One without the other produces a number nobody can act on.

On this pageApplication-path numbers set the SLO; inference-path numbers set the hardware.

DEPLOYMENT

Four Walls: nothing leaves your network

The load generator, the prompt corpus, every captured response, and the results database run on your network, on your hardware, with your load tool. No traffic to a SaaS load platform, no prompts in a vendor's cloud, no results exported for analysis. This is how we already run performance testing for regulated clients: the same tooling, the same discipline, pointed at the model instead of the EHR.

Who it is for: hospitals and research institutions, defense and government, finance, and any team running private models because the data cannot leave.

Second track: For applications built on hosted APIs, we run a cloud-API track with rate limits and cost per run modeled into the plan. Four Walls is the default we lead with; cloud API is the option.

VALIDATION AT SCALE

Verified Under Load

Speed is half the answer. Under load, AI systems fail quietly: answers truncate, timeouts return partial results, OCR skips pages when the queue backs up, a de-identifier leaks an entity it caught at low volume, an agent drops a step. Verified Under Load means we measure throughput and latency AND we inspect the outputs the system produced during the run, automatically, for exactly those failure modes. For regulated work the default is one hundred percent of a defined slice of the output; sampling is used only when the client chooses it, and it is stated as such in the report.

WHAT YOU GET
  • [·]A per-run evidence report (pass, fail, quarantine counts by failure mode).
  • [·]The load level at which output quality starts to degrade, alongside the load level at which latency does.
  • [·]The independence rule stated in writing (the checker never shares a model or a code path with the system under test).

On this pageThe count we recommend is the one at which outputs still pass, not just the one at which latency still passes.

ENGAGEMENT

How an engagement runs

01
Define
Target volume and growth, SLOs, candidate models and hardware.
02
Build
Load scenarios in your tool and AIPerf profiles; warm-cache and cold-cache variants.
03
Run
Ramp, sustained load, burst, soak (long enough to expose memory growth), and breakpoint search with fine steps near saturation.
04
Verify
Automated inspection of outputs from the runs (Verified Under Load).
05
Report
Curves, breakpoints, sizing, and the evidence report, with a recommendation the client can put in front of leadership or procurement.
DELIVERABLES

What you walk away with

IN THE REPORT
  • [·]A sizing memo with the count range and its assumptions.
  • [·]Cost per unit of throughput by option.
  • [·]The measured curves behind it.
  • [·]A procurement-ready summary.
Performance engineering on hospital and enterprise systems since 2000LoadRunner and Performance Center operated for clients continuously since 2014Test automation practice across regulated environmentsOpenText partner and reseller
FAQ

Straight answers

Can you size for a cloud GPU deployment as well as on-prem?
Yes; the method is the same, the cost model changes.
We have not chosen a model yet.
Pre-sizing is designed for that: same sample, several candidates, one comparison.
How accurate is pre-sizing?
It is a range with stated assumptions, measured on real hardware with your real workload shape. It is far closer than a vendor sheet and it tells you what to confirm in the full test.
// Start with the number you do not have

Start with the number you do not have

Bring us the system and the volume you need it to survive. We will scope the first run, tell you what it will measure, and put a date on the report.

How we prove it first
contact@proticom.com
844.PROTICOM
proticom.ai
»   REAL AI · PRODUCTION GRADE · NO HYPE