11 systems / 3 deployment classes

Compare GPU Server Prices Before You Buy AI & LLM Workstations

Filter by system class, GPU count, VRAM, workload and facility. Every result shows its price, key limitations and model compatibility.

Three-quarter supplier render of a 4U OEM multi-GPU rack server
Three-quarter supplier render of a 4U OEM multi-GPU rack server
OEM platform reference render showing the rack-server form and external service access. OEM supplier reference image.
OEM platform reference render showing the rack-server form and external service access.

All 11 systems before filtering

See the complete 4 / 5 / 2 range before narrowing it.

Every controlled product is visible here: four SME workstations, five enterprise PCIe racks and two frontier systems. Open a system directly or send a whole class into the catalogue filters.

Matching systems

Matching systems stay visible while filters change.

11 systems shown

SME workstation range

Four tower systems, from one 32GB GPU to two 96GB GPUs.

Compare form, GPU configuration, per-GPU and aggregate memory, system resources, power planning, representative fit and the first limitation before price becomes the deciding factor.

  1. Reference system 01

    Team 32

    Workstation-family reference image · appearance varies by configuration
    Workstation

    A selected-supplier RTX 5090 tower for controlled local chat, RAG, coding, speech, image and batch evaluation when the accepted workload fits one 32GB GPU.

    GPU configuration
    1 × NVIDIA GeForce RTX 5090 32GB
    GPU count
    1
    Per GPU
    32GB VRAM
    GPU memory
    32GB one GPU
    System memory
    64GB
    Primary storage
    2TB NVMe

    32GB is available on one GPU as one physical memory space.

    Power planning

    Specified for the ordered configuration and intended facility.

    Representative fit

    One team beginning a controlled local-AI deployment

    Check before selection

    32GB is one GPU's physical memory; context, cache, batching and concurrency reduce the usable model envelope.

    Model compatibility

    18 models fit without splitting the model across GPUs .

    See compatible models
  2. Reference system 02

    Company 64

    Workstation-family reference image · appearance varies by configuration
    Workstation

    A supplier-validated dual-GPU workstation for teams needing two local workers or qualified PCIe sharding, with no NVLink or pooled-memory claim.

    GPU configuration
    2 × NVIDIA GeForce RTX 5090 32GB
    GPU count
    2
    Per GPU
    32GB VRAM
    Physical total
    64GB not automatically pooled
    System memory
    128GB
    Primary storage
    2TB NVMe

    64GB is the physical total across 2 separate 32GB GPUs. It is not automatically pooled; supported software may divide a model or workload across them, subject to topology and testing.

    Power planning

    Specified for the ordered configuration and intended facility.

    Representative fit

    Two independent private AI or creative queues

    Check before selection

    64GB is aggregate capacity across two 32GB GPUs, not one universal memory pool.

    Model compatibility

    18 models fit without splitting the model across GPUs .

    See compatible models
  3. Reference system 03

    Studio 96

    Workstation-family reference image · appearance varies by configuration
    Workstation

    A selected-supplier RTX PRO 6000 Blackwell workstation for named models and professional workloads that need substantially more than 32GB on one GPU.

    GPU configuration
    1 × NVIDIA RTX PRO 6000 Blackwell 96GB
    GPU count
    1
    Per GPU
    96GB VRAM
    GPU memory
    96GB one GPU
    System memory
    128GB
    Primary storage
    2TB NVMe

    96GB is available on one GPU as one physical memory space.

    Power planning

    Specified for the ordered configuration and intended facility.

    Representative fit

    A named model or creative workload that benefits from 96GB on one GPU

    Check before selection

    The exact RTX PRO 6000 Workstation or Max-Q selection remains a pre-quote decision.

    Model compatibility

    25 models fit without splitting the model across GPUs .

    See compatible models
  4. Reference system 04

    Studio 192

    Workstation-family reference image · appearance varies by configuration
    Workstation

    A supplier-validated dual RTX PRO tower for studios and technical teams, explicitly treated as 192GB aggregate capacity rather than one pooled memory space.

    GPU configuration
    2 × NVIDIA RTX PRO 6000 Blackwell 96GB
    GPU count
    2
    Per GPU
    96GB VRAM
    Physical total
    192GB not automatically pooled
    System memory
    128GB
    Primary storage
    2TB NVMe

    192GB is the physical total across 2 separate 96GB GPUs. It is not automatically pooled; supported software may divide a model or workload across them, subject to topology and testing.

    Power planning

    Specified for the ordered configuration and intended facility.

    Representative fit

    Two independent high-memory professional GPU services

    Check before selection

    192GB is aggregate capacity across two 96GB GPUs, not one universal memory pool.

    Model compatibility

    25 models fit without splitting the model across GPUs ; 2 more need multi-GPU validation.

    See compatible models

* Aggregate VRAM is a hardware total, not a promise that every workload can use it as one memory pool. Confirm the exact model, runtime and parallelisation route.

PCIe rack range · high-memory H200 option

Five rack systems for shared workers and enterprise operation.

Choose by GPU layout, per-GPU memory, workload and facility needs. H200 adds a high-memory eight-GPU option.

  1. Reference system 01

    Value Rack 128

    GENOAX2 platform-family image showing a representative internal layout
    Rack server

    Four PCIe Gen5 GPU workers for batch, private inference and service queues, supplied as one configured enterprise system.

    GPU configuration
    4 × NVIDIA RTX 5090 AI 32GB
    GPU count
    4
    Per GPU
    32GB VRAM
    Physical total
    128GB not automatically pooled
    System memory
    Specified to order
    Primary storage
    Specified to order

    128GB is the physical total across 4 separate 32GB GPUs. It is not automatically pooled; supported software may divide a model or workload across them, subject to topology and testing.

    Power planning

    Specified for the ordered configuration and intended facility.

    Representative fit

    Four independent GPU services or queues

    Check before selection

    128GB is aggregate capacity across four 32GB GPUs.

    Model compatibility

    18 models fit without splitting the model across GPUs .

    See compatible models
  2. Reference system 02

    Value Rack 256

    GENOAX2 platform-family image showing a representative internal layout
    Rack server

    Eight PCIe Gen5 GPU workers for dense batch and service throughput. Selection depends on sustained utilisation, final topology, GPU warranty and facility acceptance.

    GPU configuration
    8 × NVIDIA RTX 5090 AI 32GB
    GPU count
    8
    Per GPU
    32GB VRAM
    Physical total
    256GB not automatically pooled
    System memory
    Specified to order
    Primary storage
    Specified to order

    256GB is the physical total across 8 separate 32GB GPUs. It is not automatically pooled; supported software may divide a model or workload across them, subject to topology and testing.

    Power planning

    Specified for the ordered configuration and intended facility.

    Representative fit

    Eight independent services or parallel batch queues

    Check before selection

    256GB is aggregate capacity across eight 32GB GPUs.

    Model compatibility

    18 models fit without splitting the model across GPUs ; 5 more need multi-GPU validation.

    See compatible models
  3. Reference system 03

    Enterprise 384

    GENOAX2 platform-family image showing a representative internal layout
    Rack server

    Four RTX PRO 6000 Blackwell Server Edition GPU workers for controlled enterprise services, with platform, cooling and support brought together.

    GPU configuration
    4 × NVIDIA RTX PRO 6000 Blackwell 96GB
    GPU count
    4
    Per GPU
    96GB VRAM
    Physical total
    384GB not automatically pooled
    System memory
    Specified to order
    Primary storage
    Specified to order

    384GB is the physical total across 4 separate 96GB GPUs. It is not automatically pooled; supported software may divide a model or workload across them, subject to topology and testing.

    Power planning

    Specified for the ordered configuration and intended facility.

    Representative fit

    Several 96GB private inference endpoints

    Check before selection

    384GB is aggregate capacity across four 96GB GPUs.

    Model compatibility

    25 models fit without splitting the model across GPUs ; 5 more need multi-GPU validation.

    See compatible models
  4. Reference system 04

    Enterprise 768

    GENOAX2 platform-family image showing a representative internal layout
    Rack server

    Eight RTX PRO 6000 Blackwell Server Edition GPU workers for large private platforms where independent services or tested sharding justify the facility requirements.

    GPU configuration
    8 × NVIDIA RTX PRO 6000 Blackwell 96GB
    GPU count
    8
    Per GPU
    96GB VRAM
    Physical total
    768GB not automatically pooled
    System memory
    Specified to order
    Primary storage
    Specified to order

    768GB is the physical total across 8 separate 96GB GPUs. It is not automatically pooled; supported software may divide a model or workload across them, subject to topology and testing.

    Power planning

    Specified for the ordered configuration and intended facility.

    Representative fit

    Eight high-memory endpoints or batch workers

    Check before selection

    768GB is aggregate capacity across eight 96GB GPUs.

    Model compatibility

    25 models fit without splitting the model across GPUs ; 6 more need multi-GPU validation.

    See compatible models
  5. Reference system 05

    H200 1.1TB

    H200-capable GENOAX2 platform-family image showing a representative eight-GPU topology and internal layout
    Rack server

    A quote-only eight-GPU PCIe H200 NVL system for qualified large-model work. Kimi K3 and other large models require an exact runtime profile and acceptance test.

    GPU configuration
    8 × NVIDIA H200 NVL 141GB
    GPU count
    8
    Per GPU
    141GB VRAM
    Physical total
    1128GB not automatically pooled
    System memory
    1024GB
    Primary storage
    11.68TB NVMe

    1,128GB is the physical total across eight H200 GPUs. It is not automatically one pooled memory space; supported parallel operation depends on the selected topology and runtime.

    Power planning

    Specified for the ordered configuration and intended facility.

    Representative fit

    A named large model with a reproduced eight-GPU profile

    Check before selection

    1,128GB aggregate HBM is below the current 1.5609TB official Kimi K3 weight-file total.

    Model compatibility

    25 models fit without splitting the model across GPUs ; 6 more need multi-GPU validation.

    See compatible models

* Aggregate VRAM is a hardware total, not a promise that every workload can use it as one memory pool. Confirm the exact model, runtime and parallelisation route.

Frontier data-centre systems

Native HGX B300 and full-rack GB300 NVL72.

Both are configured to order for specialist data-centre projects. Compare GPU memory with the intended models, software, workload and performance requirement.

  1. Reference system 01

    Frontier Native 2.3TB

    SYS-822GS-NB3RT B300 platform reference image
    Rack server

    A quote-only HGX B300 route using eight 288GB SXM GPUs and NVLink/NVSwitch architecture, subject to exact OEM configuration, software support and facility design.

    GPU configuration
    8 × NVIDIA B300 288GB
    GPU count
    8
    Per GPU
    288GB VRAM
    Physical total
    2304GB not automatically pooled
    System memory
    Specified to order
    Primary storage
    Specified to order

    2,304GB is distributed across eight B300 GPUs connected by HGX NVLink and NVSwitch. The fabric enables high-bandwidth multi-GPU work, but software support and the exact workload still determine usable capacity.

    Power planning

    Specified for the ordered configuration and intended facility.

    Representative fit

    A named model with a validated HGX B300 runtime

    Check before selection

    2.304TB HBM is a calculated capacity envelope, not proof that the official Kimi K3 checkpoint loads or serves one-million-token context.

    Model compatibility

    30 models fit without splitting the model across GPUs ; 3 more need multi-GPU validation.

    See compatible models
  2. Reference system 02

    Frontier Rack 20TB

    GB300 NVL72 rack architecture reference image
    Rack server

    A quote-only GB300 NVL72 route with 72 Blackwell GPUs and 36 Grace CPUs. It is a complete rack-scale platform, not an ordinary GPU server.

    GPU configuration
    72 × NVIDIA GB300 288GB
    GPU count
    72
    Per GPU
    288GB VRAM
    Physical total
    20736GB not automatically pooled
    System memory
    Specified to order
    Primary storage
    Specified to order

    20,736GB is distributed across a rack-scale NVLink and NVSwitch fabric. It is specialist shared infrastructure, not one ordinary GPU memory space.

    Power planning

    Up to 142kW full-rack power

    Representative fit

    A validated rack-scale model or training architecture

    Check before selection

    This is a complete 72-GPU liquid-cooled rack, not an the selected OEM 4U product or ordinary office delivery.

    Model compatibility

    30 models fit without splitting the model across GPUs ; 3 more need multi-GPU validation.

    See compatible models

* Aggregate VRAM is a hardware total, not a promise that every workload can use it as one memory pool. Confirm the exact model, runtime and parallelisation route.

Included baseline

Specified, configured, tested and handed over.

The listed price covers the physical build, initial software configuration and documented remote onboarding. It is not a promise of open-ended implementation or a managed service.

  • Itemised hardware specification
  • Ubuntu and validated NVIDIA software pairing
  • One agreed inference runtime and browser interface
  • Model source and licence record
  • 24-hour minimum burn-in and test evidence
  • Credential and security-baseline handover
  • Remote onboarding
  • 30-day configuration-defect support

Do not choose from a card

Give the workload a chance to reject the hardware.

The right outcome may be a smaller machine, cloud service, colocation, hybrid service - or no purchase.

Tell us what you need

Decision check

GPU Server Price: Compare Like-for-Like Prices

For GPU server price, compare the complete configured system rather than the accelerator headline alone.

A useful GPU server price comparison should show the specification, evidence status, facility needs and quotation boundary. Include AI requirements, workstation form and LLM requirements where those factors change the decision.