Enterprise accelerator rack · eight H200 NVL GPUs
H200 1.1TB
A quote-only 1.1TB aggregate H200 NVL platform for qualified large-model work.
A quote-only eight-GPU PCIe H200 NVL system for qualified large-model work. Kimi K3 and other large models require an exact runtime profile and acceptance test.
Current quotation and acceptance plan required.
Price, exact parts, compatibility, warranty, delivery and workload results are indicative until confirmed in a written quotation and acceptance record.
Key buying facts
H200 1.1TB at a glance
- Commercial status
- Quote only Indicative budget £349,000 ex VAT current supplier quotation required
- GPU configuration
- 8 × NVIDIA H200 NVL 141GB 8 GPUs
- Per-GPU memory
- 141GB physical VRAM per GPU
- Physical GPU total
- 1128GB across 8 GPUs · not automatically pooled
- Physical class
- Rack server Eight-GPU H200 NVL PCIe rack platform
How the memory works: 1,128GB is the physical total across eight H200 GPUs. It is not automatically one pooled memory space; the exact NVLink pairing, PCIe topology and supported parallel runtime must be confirmed in the quotation.
Power planning · Rack server
At least 4.8kW GPU nameplate; exact full-system input pending
Representative intended fit
Enterprise inference or fine-tuning evaluation
Buyer fit
Start with the reason to own it.
A qualified enterprise route when H200 software maturity or availability is preferred over newer Blackwell platforms.
Commercial status
Quote only
Final supplier quote required
Indicative budget £349,000 ex VAT · final quotation required
Includes workload sizing, the configured system, AI software stack, security baseline, burn-in, agreed workload testing, documentation, remote onboarding and 30-day configuration-defect support.
Package figures exclude VAT and cover the stated GPU RIGS deployment scope. Every final order requires a written quotation confirming the specification, availability, delivery, warranty and workload acceptance plan.
A credible fit
- A named large model with a reproduced eight-GPU profile
- Enterprise inference or fine-tuning evaluation
- A buyer with data-centre power, cooling and operations
Choose another route when
- A PCIe RTX PRO worker platform meets demand
- The purchase depends on unverified speed or concurrency
- The exact supplier and support route is unresolved
System evidence
What is confirmed, calculated and still unknown.
Vendor-documented values, calculations, assumptions and unknowns remain separate. A capability is marked as measured only when a reproducible benchmark exists.
- Commercial status
- quote-only
- Reviewed
- Review again
- Open specification items
- 9
01 / model evidence
Exact checkpoint profiles
0 measured-supported profiles. A named model appears as supported only after the exact checkpoint, runtime, deployment and limits have a reproducible benchmark record.
Current boundary
The current 1,560,936,091,448-byte Kimi K3 repository weight total exceeds this system's 1,128GB aggregate HBM. An aggressively quantised conversion has not been validated; no converted artifact, runtime, context or benchmark has been recorded.
This does not mean that every model is unsupported. It means untested memory arithmetic or a generic GPU claim is not treated as a model match.
02 / workload acceptance
Beyond language models
Each workload needs its own quality, latency, throughput, stability and resource test. The entries below define what would be measured; they are not performance claims.
01
Retrieval-augmented generation
Retrieval quality, citation support, permissions, refusal behaviour and latency against a versioned corpus and question set; document count alone is not a hardware metric.
Acceptance definition only. No performance, quality, capacity or suitability result has been measured for this system.
02
Software-engineering assistance
Task correctness, test pass rate, unsafe-change rate, reviewer effort and useful response time on a versioned repository evaluation set.
Acceptance definition only. No performance, quality, capacity or suitability result has been measured for this system.
03
Speech to text
Word error rate, real-time factor and failure rate on a versioned, representative audio set.
Acceptance definition only. No performance, quality, capacity or suitability result has been measured for this system.
04
Embedding
Vectors per second, query latency and retrieval-quality metric on a versioned corpus with the exact embedding checkpoint and dimensions.
Acceptance definition only. No performance, quality, capacity or suitability result has been measured for this system.
05
Vision and OCR
Field accuracy or character error rate plus latency on a versioned, permission-safe image and document set.
Acceptance definition only. No performance, quality, capacity or suitability result has been measured for this system.
06
Image generation
Latency percentiles, images per second, peak memory, stability and accepted-output rate at exact checkpoint, resolution, steps, sampler, CFG and batch.
Acceptance definition only. No performance, quality, capacity or suitability result has been measured for this system.
07
Video generation
Clip latency, clips per hour, peak memory, stability and accepted-output rate at exact checkpoint, dimensions, frames, frame rate, duration, steps and batch.
Acceptance definition only. No performance, quality, capacity or suitability result has been measured for this system.
08
GPU rendering
Frame latency, throughput, errors, wall power and temperature for an immutable scene using the exact renderer, version, resolution and sample count.
Acceptance definition only. No performance, quality, capacity or suitability result has been measured for this system.
09
Language-model inference
Quality pass rate, TTFT, decode and aggregate throughput, latency percentiles, errors, memory, power and temperature at disclosed context, batch and concurrency.
Acceptance definition only. No performance, quality, capacity or suitability result has been measured for this system.
10
Fine-tuning
Completed steps, time, peak memory, loss and held-out evaluation result with exact base checkpoint, method, trainable parameters, dataset, sequence length and batch.
Acceptance definition only. No performance, quality, capacity or suitability result has been measured for this system.
03 / configuration
Specification with its certainty attached
The current sources show the basis for each value. The final quotation names the exact supplier, parts, warranty and bill of materials for the ordered system.
| Field | Public value | Evidence state | Qualification |
|---|---|---|---|
| Platform class | Eight-GPU H200 NVL PCIe rack platform | Vendor documented | Documented by the component or platform vendor; not independently measured by GPU RIGS. This is a configurable PCIe H200 NVL platform, not an H200 SXM/HGX system. Written topology and fulfilment confirmation remain required. |
| Chassis class | Rack chassis pending written supplier topology | Vendor documented | Documented by the component or platform vendor; not independently measured by GPU RIGS. The public configurator proves a selectable platform family, not an accepted serialised GPU RIGS bill of materials. |
| GPU route | 8 × NVIDIA H200 NVL 141GB | Vendor documented | Documented by the component or platform vendor; not independently measured by GPU RIGS. Exact board part numbers and serials belong in the final bill of materials. |
| Physical GPU memory total | 1128GB across 8 GPUs | Calculated | Calculated from disclosed inputs; not a measured system result. 1,128GB is the physical total across eight H200 GPUs. It is not automatically one pooled memory space; the exact NVLink pairing, PCIe topology and supported parallel runtime must be confirmed in the quotation. |
| Per-GPU memory | 141GB | Vendor documented | Documented component capacity and the safer first sizing boundary before any supported multi-GPU test. |
| GPU interconnect | Exact H200 NVL card pairing and eight-GPU topology pending supplier validation | Unknown | Not yet evidenced for the ordered system and must be resolved before acceptance. Aggregate memory alone does not establish an eight-GPU topology. The supplier quote must name NVLink pairing, PCIe layout and any supported collective path. |
| System topology | Exact H200 NVL pair, PCIe lane and collective topology pending | Unknown | Not yet evidenced for the ordered system and must be resolved before acceptance. The configurator selects eight passive PCIe 5.0 x16 H200 NVL cards but exposes no H200-specific NVLink option. Do not infer an HGX NVSwitch fabric or a working collective topology from aggregate HBM capacity. |
| CPU | 2 × AMD EPYC 9454 48-core | Vendor documented | Documented by the component or platform vendor; not independently measured by GPU RIGS. Selected public reference build; final CPU choice remains subject to the accepted workload and supplier quotation. |
| System memory | 1,024GB DDR5 ECC RDIMM (16 × 64GB) | Vendor documented | Documented by the component or platform vendor; not independently measured by GPU RIGS. Selected public reference build; DIMM population and validated memory speed remain quotation items. |
| Primary storage | 11.68TB raw NVMe (2 × 3.84TB U.2 plus 2 × 2TB M.2) | Vendor documented | Documented by the component or platform vendor; not independently measured by GPU RIGS. Raw device capacity before formatting, RAID, endurance policy and usable-capacity decisions. |
| Network | 200Gb ConnectX-6 plus management/base networking | Vendor documented | Documented by the component or platform vendor; not independently measured by GPU RIGS. Cables, switch compatibility, fabric mode and measured throughput remain outside the public subtotal. |
| Cooling | Forced-air cooling design; exact implementation pending | Assumption | A planning assumption that must not be treated as a confirmed specification. The chassis uses high-speed hot-swap fans, but a completed eight-H200 thermal and acoustic acceptance result is not published. |
| Commercial status | Quote only | Unknown | No public numeric offer is made. Exact hardware, facility, delivery, warranty and support scope require a written quotation. |
04 / site and facility
The room stays inside the system boundary
Nameplate power, supplier cooling descriptions and aggregate GPU TDP are not a site approval. The exact build and customer facility need written acceptance.
- Input power
- At least 4.8kW GPU nameplate; exact full-system input pending Calculated Calculated from disclosed inputs; not a measured system result. Calculated from 8 × 600W GPU TDP and excludes CPU, memory, storage, fans and conversion losses.
- Measured heat output
- Measured wall power and facility heat-rejection requirement pending reference build Unknown Not yet evidenced for the ordered system and must be resolved before acceptance.
- Cooling route
- Forced-air cooling design; exact implementation pending Assumption A planning assumption that must not be treated as a confirmed specification. The chassis uses high-speed hot-swap fans, but a completed eight-H200 thermal and acoustic acceptance result is not published.
- Dimensions and handling
- 940 × 440 × 176.5mm supplier-listed chassis dimensions Vendor documented Documented by the component or platform vendor; not independently measured by GPU RIGS. Final packed dimensions, weight, rack depth and service-clearance requirements remain to be confirmed.
- Facility acceptance
- Rack, power, airflow, heat rejection, noise, UPS and service access required Unknown The customer site has not been surveyed. The final ordered configuration and qualified facility design control this decision.
Pre-order site checks
- A suitable 19-inch rack, rail depth, handling route and secure operating location
- A qualified electrical design based on the final PSU population and measured load
- Cooling and heat-rejection capacity for sustained accelerator operation
- Appropriate switching, cabling, remote management and network segmentation
- A named operational owner or contracted support route
05 / delivery and support
A handover boundary, not an implied managed service
The final quotation must name who owns procurement, facility work, acceptance, warranty, remote support, recurring operations and any on-site service.
- Hardware warranty
- Three-year parts warranty listed; back-to-back GPU RIGS support route pending Vendor documented Documented by the component or platform vendor; not independently measured by GPU RIGS. Exact response times, advance replacement, freight liability, exclusions and UK service process require written supplier terms.
- Availability and lead time
- Build-to-order availability and lead time pending written supplier confirmation Unknown Not yet evidenced for the ordered system and must be resolved before acceptance. This record does not mean the product is stocked or orderable.
- GPU RIGS support baseline
- Remote onboarding and 30-day configuration-defect support Assumption The written quotation and contract confirm this service boundary. It is not a 24-hour managed service or an implied on-site warranty.
Standard service boundary
- Documented workload and site-fit review
- Confirmed bill of materials before procurement
- Configuration, burn-in and agreed smoke-test evidence
- Asset schedule, admin notes and user quick-start material
- Collection or the quoted kerbside or pallet-delivery route
- Remote onboarding and 30-day configuration-defect support
Separate scope or customer responsibility
- Building electrical work, rack, UPS, cooling or structured cabling
- Nationwide on-site installation unless separately quoted
- Migration of customer data, every integration or every application
- Continuous managed operations, security monitoring or a 24-hour support agreement
- Third-party model, API, marketplace or software charges
- A compliance certificate, performance guarantee or income guarantee
06 / sources and review status
Sources and unresolved items
Review dates show evidence freshness, not future availability. A current quote, serialised bill of materials and accepted test record control the order.
- Commercial state
- quote-only
- Reviewed
- 2026-07-28
- Review again
- 2026-08-28
Unresolved before acceptance
- GPU interconnect: Exact H200 NVL card pairing and eight-GPU topology pending supplier validation
- System topology: Exact H200 NVL pair, PCIe lane and collective topology pending
- Cooling: Forced-air cooling design; exact implementation pending
- Commercial status: Quote only
- Measured heat output: Measured wall power and facility heat-rejection requirement pending reference build
- Cooling route: Forced-air cooling design; exact implementation pending
- Facility acceptance: Rack, power, airflow, heat rejection, noise, UPS and service access required
- Availability and lead time: Build-to-order availability and lead time pending written supplier confirmation
- GPU RIGS support baseline: Remote onboarding and 30-day configuration-defect support
-
Product record
Current product specification and commercial status
GPU RIGS
Retrieved 2026-07-28 · reviewed 2026-07-28
The current product record defines the specification and commercial status shown on this page.
-
Retrieved 2026-07-28 · reviewed 2026-07-28
Use the linked source to confirm the dated component, platform or checkpoint facts stated above.
-
Supplier reference
Dated supplier configuration evidence
System supplier
Retrieved 2026-07-28 · reviewed 2026-07-28
Dated supplier evidence. Configuration, stock, price and warranty require confirmation in the final quotation.
-
Supplier reference
Dated supplier configuration evidence
System supplier
Retrieved 2026-07-28 · reviewed 2026-07-28
Dated supplier evidence. Configuration, stock, price and warranty require confirmation in the final quotation.
-
Market comparison
Dated market comparison
Market reference
Retrieved 2026-07-28 · reviewed 2026-07-28
Dated market comparison only; not an exact quotation or proof of product equivalence.
Known limitations
- 1,128GB aggregate HBM is below the current 1.5609TB official Kimi K3 weight-file total.
- Kimi K3 is only an aggressive-quantisation candidate on this route until an exact converted checkpoint is tested.
- No native load, full context, tokens-per-second, concurrency or quality result is claimed.
- 1,128GB aggregate HBM is below the current 1.5609TB official Kimi K3 weight-file total.
- Kimi K3 is only an aggressive-quantisation candidate on this route until an exact converted checkpoint is tested.
- No native load, full context, tokens-per-second, concurrency or quality result is claimed.
- The current £349,000 ex-VAT planning figure is indicative; the system remains quote-only until the supplier BOM, topology, landed cost and support route are accepted.
- Kimi K3 is a calculated capacity candidate, not a supported or benchmarked claim.
- Aggregate GPU memory is not one universal memory pool.
- No model, speed, context, concurrency or multi-GPU efficiency promise exists without a controlled fit profile.
- Rack, electrical work, UPS, cooling, cabling, colocation and on-site installation require a separate scope.
Do not buy this route when
- The requirement is the current official Kimi K3 checkpoint at native format.
- The buyer cannot accept quote-only topology, facility and support discovery.
- HGX B300 or hosted frontier capacity offers a more appropriate memory/interconnect route.
Inspect the evidence behind H200 1.1TB.
The system view shows the selected hardware reference where available. Technical diagrams explain memory, data boundaries, queues, power and handover. The final bill of materials and workload test confirm the ordered system.
Software and security
The usable product is more than the chassis.
The final stack stays deliberately small. Versions, licences, access and recurring ownership are recorded so the customer is not left with an opaque collection of containers.
Software baseline
- Ubuntu LTS on a recorded operating-system version
- NVIDIA driver, CUDA components and container support validated for the ordered hardware
- Docker Engine and NVIDIA Container Toolkit
- One primary model server selected from Ollama, vLLM, SGLang, TensorRT-LLM or llama.cpp for the accepted workload
- Open WebUI or another reviewed browser interface
- Named authentication, TLS and reverse-proxy approach
- GPU, node and service monitoring with an agreed log-retention period
- Pinned versions, a software bill of materials and a model source and licence record
Security ownership
- GPU RIGS baseline
- Initial operating-system state, named administrator handover, host firewall baseline, agreed access route, secrets transfer and documented update state.
- Customer or contracted operator
- User lifecycle, network and VPN policy, backups, monitoring review, patch approval, incident response, data governance and lawful use.
- Shared before acceptance
- Model and software licence checks, retention and logging choices, recovery test, acceptance criteria and a named owner for every recurring task.
Local infrastructure can reduce disclosure to external AI APIs. It does not automatically make the service secure, accurate, confidential or UK GDPR compliant.
Testing and acceptance
The evidence pack is part of the machine.
Performance is confirmed against the accepted workload, exact build and disclosed test conditions. No throughput, latency or quality figure is claimed before that test.
- 01
Record the final bill of materials, serial numbers and firmware versions.
- 02
Run at least 24 hours of GPU, CPU, memory and storage stress testing.
- 03
Capture temperature, fan, error, health and wall-power evidence under the agreed load.
- 04
Test cold boot, restart and the available remote-management route.
- 05
Check drive health, network throughput and the container and GPU runtime.
- 06
Run model and workload smoke tests against the written acceptance set.
- 07
Record any measured speed only with the model, quantisation, context, concurrency and runtime disclosed.
- 08
Test an agreed fault or recovery route and provide the resulting handover record.
What is not claimed today
No fixed users, tokens per second, latency, model size, accuracy, availability, savings or marketplace contribution is stated without the missing configuration and test conditions.
See how systems are testedCommercial reality
Tax and spare capacity are supporting questions.
Neither belongs in a guaranteed saving or payback claim. The purchase must stand on the accepted workload, control case and operating plan.
Finance, VAT and capital allowances
- The displayed figure is for a complete GPU RIGS private AI deployment, not unconfigured hardware. The final written quotation confirms the exact specification, delivery, warranty and accepted workload scope.
- Prices exclude VAT. VAT recovery depends on the buyer, its taxable activities and the normal evidence rules.
- Qualifying equipment may be plant and machinery for capital-allowance purposes. The buyer's accountant decides eligibility and timing.
- A third-party lease or hire-purchase route may be explored after partner validation and credit approval. GPU RIGS is not presented as a lender.
- There is no generic capital-gains advantage and no automatic research and development relief because the equipment supports AI.
Optional idle capacity
- Marketplace mode is off by default and is excluded from the purchase case.
- A separate environment, no customer data mounts, network controls and a local kill switch would be required.
- The customer, insurer, supplier warranty and marketplace terms must permit the intended use.
- Vast.ai, Render, Golem and direct batch work do not guarantee acceptance, demand, rate or income.
- Any dated estimate must identify its source date; a pilot must report achieved utilisation, fees, electricity, cooling, faults and operator time.
Limits and alternatives
A good specification leaves room for “no”.
The fit check can recommend a smaller system, hosted service, hybrid route, custom build or no purchase. That is preferable to making this package carry a workload it has not proved.
Package boundaries
- Aggregate GPU memory is not one universal memory pool.
- No model, speed, context, concurrency or multi-GPU efficiency promise exists without a controlled fit profile.
- Rack, electrical work, UPS, cooling, cabling, colocation and on-site installation require a separate scope.
- 1,128GB aggregate HBM is below the current 1.5609TB official Kimi K3 weight-file total.
- Kimi K3 is only an aggressive-quantisation candidate on this route until an exact converted checkpoint is tested.
- No native load, full context, tokens-per-second, concurrency or quality result is claimed.
- The current £349,000 ex-VAT planning figure is indicative; the system remains quote-only until the supplier BOM, topology, landed cost and support route are accepted.
- Kimi K3 is a calculated capacity candidate, not a supported or benchmarked claim.
- Aggregate GPU memory is not one universal memory pool.
- No model, speed, context, concurrency or multi-GPU efficiency promise exists without a controlled fit profile.
- Rack, electrical work, UPS, cooling, cabling, colocation and on-site installation require a separate scope.
Choose the smaller system when it meets the accepted workload and resilience requirement.
Consider this route Custom specificationUse when storage, network, resilience, colocation or service design differs from the reference route.
Consider this route Hosted or hybrid AIMay be preferable for bursty demand or where the organisation cannot own the facility and operating burden.
Consider this routeQuestions answered
H200 1.1TB questions that affect the order
The written quotation and acceptance plan confirm any price, specification, warranty or workload assumption discussed below.
Does H200 1.1TB provide 1128GB as one memory pool?
Only a single-GPU system provides its stated GPU memory on one card. Multi-GPU totals are aggregate physical capacity; usable sharding depends on the exact model, runtime and topology.
Which models are supported?
Only models with a reproducible fit profile are listed as supported. Capacity candidates are labelled separately; a memory calculation alone is not a support claim.
Is the displayed price final?
No. An indicative price is a transparent planning value based on dated inputs. Quote-only platforms and the final order require a current supplier quote, accepted bill of materials and delivery scope.
Can spare capacity earn income?
It may be evaluated as an optional isolated secondary use. Marketplace acceptance, utilisation, rates, fees, energy and income can change and are never guaranteed or included in the purchase case.
Prepare the next decision
Specify the work before the parts.
Record the workload, data boundary, users, site and acceptance test. No confidential documents or credentials are needed for the first brief.
Package figures exclude VAT and cover the stated GPU RIGS deployment scope. Every final order requires a written quotation confirming the specification, availability, delivery, warranty and workload acceptance plan.