First-pass technical fit

Private AI server sizing tool.

Choose a memory band, active demand and location. The result narrows what to test next; it is not a benchmark or promise that a model will fit.

Diagram combining model weights, context, cache and active requests into a memory headroom check
Diagram combining model weights, context, cache and active requests into a memory headroom check
Memory fit depends on the workload, context, cache and simultaneous demand, then needs a representative test. GPU Servers technical illustration.
Memory fit depends on the workload, context, cache and simultaneous demand, then needs a representative test.
Use the exact model, quantisation, context and runtime to verify this.
Use active requests, not named staff.

Next hardware evaluation

-

  • Still requiredExact workload test
  • Still requiredPower and site pre-flight
  • Still requiredSupplier and warranty check

Calculation record

See what sits behind the result.

The controls are deliberately editable. Record the values used, the date and the source before relying on an output in a buying decision.

01

Method

  • The first branch is the likely minimum memory required on one GPU.
  • Active generations alter queue and cache demand; named users do not.
  • Location changes the viable acoustic, power, cooling and service boundary.
02

Worked reading

  • A 24GB target with modest concurrent demand may justify testing a compact system.
  • A model needing more than 48GB on one GPU points to a 96GB-class test, not two unrelated 24GB cards.
  • Any office result still needs a noise and facilities check.
03

Do not infer

  • The tool does not prove that a model, context or runtime will fit.
  • Aggregate VRAM is not automatically one shared memory pool.
  • Only a representative benchmark can establish useful latency, throughput and quality.

Why the tool stops at a hardware class

Model weights, context, KV cache, batching, runtime overhead and inter-GPU behaviour change the result. User count alone cannot produce a safe server recommendation. The result narrows the hardware that should be tested next.

Learn how local LLM memory works.

Management access belongs in the specification.

A useful handover records software versions, credentials, recovery, monitoring and update ownership alongside the physical server.

Diagram showing an approved request, a local service, an approved store and a policy-controlled data path
Diagram showing an approved request, a local service, an approved store and a policy-controlled data path
A private deployment starts with the permitted data path, access policy and logging boundary. GPU Servers technical illustration.
A private deployment starts with the permitted data path, access policy and logging boundary. GPU Servers technical illustration.
GPU server remote management dashboard with system status, access logs and sensor monitoring panels
GPU server remote management dashboard with system status, access logs and sensor monitoring panels
Supplier screenshot of the platform management interface. The final management features and access policy depend on the ordered system. OEM supplier reference image.
Supplier screenshot of the platform management interface. The final management features and access policy depend on the ordered system. OEM supplier reference image.
Diagram showing approved documents moving through a searchable index to an answer with a source citation
Diagram showing approved documents moving through a searchable index to an answer with a source citation
A retrieval workflow should connect each useful answer to approved source material and defined refusal behaviour. GPU Servers technical illustration.
A retrieval workflow should connect each useful answer to approved source material and defined refusal behaviour. GPU Servers technical illustration.
Diagram of an evidence pack containing an asset schedule, burn-in record, health readings, workload test and admin guide
Diagram of an evidence pack containing an asset schedule, burn-in record, health readings, workload test and admin guide
A credible handover records the supplied assets, checks, operating evidence, instructions and unresolved items. GPU Servers technical illustration.
A credible handover records the supplied assets, checks, operating evidence, instructions and unresolved items. GPU Servers technical illustration.