OEM enterprise rack · eight 96GB GPUs
Buy the Enterprise 768 8-GPU RTX 6000 Server
Eight high-memory Server Edition workers at the top of the PCIe range.
Eight RTX PRO 6000 Blackwell Server Edition GPU workers for large private platforms where independent services or tested sharding justify the facility requirements.
Key buying facts
Enterprise 768 at a glance
- Price
- £189,999 guide price · complete configured system
- GPU configuration
- 8 × NVIDIA RTX PRO 6000 Blackwell 96GB 8 GPUs
- Per-GPU memory
- 96GB physical VRAM per GPU
- Physical GPU total
- 768GB across 8 GPUs · not automatically pooled
- Physical class
- Rack server GENOAX2 PCIe Gen5 4U enterprise rack platform
How the memory works: 768GB is the physical total across 8 separate 96GB GPUs. It is not automatically pooled; supported software may divide a model or workload across them, subject to topology and testing.
Power planning · Rack server
Specified for the ordered configuration and intended facility.
Representative intended fit
A validated PCIe multi-GPU runtime
Buyer fit
Start with the reason to own it.
The highest PCIe worker tier before moving to tightly coupled frontier platforms.
Price
£189,999
guide price · complete configured system
Includes workload sizing, the configured system, AI software stack, security baseline, burn-in, agreed workload testing, documentation, remote onboarding and 30-day configuration-defect support.
Guide prices cover the system and service scope described on this page. The final total is confirmed before purchase.
A credible fit
- Eight high-memory endpoints or batch workers
- A validated PCIe multi-GPU runtime
- A data-centre-grade operating and support route
Choose another route when
- Four workers meet demand
- The workload requires an HGX fabric
- Power, cooling or support is outstanding
Model compatibility
Models that fit Enterprise 768's GPU memory
96GB is available per GPU, with 768GB physically installed across 8 GPUs.
The results below compare that hardware with each model's GPU-memory requirement. They do not predict speed, maximum context, image size or concurrent users.
All 33 models in our current model guide are included below. Another version needs its own exact checkpoint, licence, runtime and memory requirement before it can be matched to hardware.
25 models fit without splitting the model across GPUs
6 models may fit when supported software divides the model across GPUs
Compare every modelLarger-system opportunities
Choose a larger server for 2 additional models
If these models are part of your plan, the routes below show the smallest larger system that best meets their GPU-memory needs. Compare it with Enterprise 768, then size the final configuration around your workload.
Multi-GPU opportunity
Frontier Native 2.3TB
The first larger system with enough installed GPU memory when supported software divides these models across its GPUs.
- Kimi K3 At least 1.561TB aggregate GPU memory for the released weight files
- GLM-5 1536GB aggregate across at least 8 GPUs
We confirm software support, GPU layout, context and performance before purchase.
Qwen
Qwen3.6 35B-A3B
A 36-billion-parameter multimodal mixture-of-experts model with about 3 billion active parameters, a 262,144-token default context and a strong emphasis on agentic coding.
- Minimum GPU memory
- 80GB-class GPU for the named BF16 weights at reduced context
- Licence
- Apache License 2.0
- Useful for
- Language & reasoning, Coding & agents, Vision & OCR
OpenAI
GPT-OSS 120B
OpenAI's 117B-total, 5.1B-active open-weight reasoning and agentic model, released with native MXFP4 MoE weights.
- Minimum GPU memory
- One 80GB GPU
- Licence
- Apache License 2.0
- Useful for
- Language & reasoning, Coding & agents
Gemma 3 27B
Google's 27B instruction-tuned multimodal model for text and image input, with a 128K context and broad language coverage.
- Minimum GPU memory
- 64GB GPU for BF16 at a bounded context
- Licence
- Gemma Terms of Use
- Useful for
- Language & reasoning, Vision & OCR
DeepSeek
DeepSeek-OCR 2
A 3B-class image-to-text model for document optical character recognition and layout-aware text extraction.
- Minimum GPU memory
- 12GB GPU planning floor
- Licence
- Apache License 2.0
- Useful for
- Vision & OCR
Qwen
Qwen3-Embedding 8B
The largest Qwen3 embedding checkpoint, designed for multilingual dense retrieval, classification, clustering and text matching.
- Minimum GPU memory
- 20GB GPU planning floor
- Licence
- Apache License 2.0
- Useful for
- Embeddings & RAG
BAAI
BGE-M3
A multilingual embedding model that supports dense, sparse and multi-vector retrieval with 1,024 dimensions and inputs up to 8,192 tokens.
- Minimum GPU memory
- 4GB GPU planning floor, or CPU for light use
- Licence
- MIT License
- Useful for
- Embeddings & RAG
Qwen
Qwen3-ASR 1.7B
A compact automatic speech recognition checkpoint released in July 2026 for multilingual transcription and audio understanding.
- Minimum GPU memory
- 8GB GPU planning floor
- Licence
- Apache License 2.0
- Useful for
- Speech & audio
OpenAI
Whisper Large V3
OpenAI's 1.55B-parameter multilingual speech-recognition and translation checkpoint, widely supported across transcription runtimes.
- Minimum GPU memory
- 6GB GPU planning floor for an optimised inference runtime
- Licence
- Apache License 2.0
- Useful for
- Speech & audio
Mistral AI
Voxtral Mini 4B Realtime
A 4B-class real-time automatic speech recognition model released by Mistral AI for low-latency streaming transcription.
- Minimum GPU memory
- 24GB GPU planning floor
- Licence
- Apache License 2.0
- Useful for
- Speech & audio
Black Forest Labs
FLUX.2 Klein 4B
A compact rectified-flow model for text-to-image, image editing and multi-reference work, released under Apache 2.0.
- Minimum GPU memory
- About 13GB VRAM
- Licence
- Apache License 2.0
- Useful for
- Image generation
Qwen
Qwen-Image 2512
The December 2025 Qwen-Image update for text-to-image generation, with improved realism, natural detail and text rendering.
- Minimum GPU memory
- 64GB GPU as a tight full-weight floor
- Licence
- Apache License 2.0
- Useful for
- Image generation
Wan AI
Wan2.2 TI2V 5B
A 5B text-and-image-to-video model that supports 720p generation and an official single-GPU offload route.
- Minimum GPU memory
- 24GB VRAM with documented CPU offload settings
- Licence
- Apache License 2.0
- Useful for
- Video generation
Tencent
HunyuanVideo 1.5
An 8.3B text-to-video and image-to-video model with 480p and 720p checkpoints, optional super-resolution and distilled workflows.
- Minimum GPU memory
- 14GB VRAM with model offloading
- Licence
- Tencent Hunyuan Community Licence
- Useful for
- Video generation
Tencent
Hunyuan3D 2.1
Tencent's image-to-3D generation pipeline for shape and texture creation, with open model weights and a model-specific community licence.
- Minimum GPU memory
- 24GB single-GPU planning floor
- Licence
- Tencent Hunyuan Community Licence
- Useful for
- 3D generation
Meta
Llama Guard 4
A 12B multimodal safeguard model for classifying text and image prompts and responses against Meta's hazard taxonomy.
- Minimum GPU memory
- 28GB GPU planning floor
- Licence
- Llama 4 Community Licence
- Useful for
- Safety & moderation, Vision & OCR
Stability AI
Stable Diffusion 3.5 Large
An 8B-class text-to-image model in the Stable Diffusion 3.5 family, with a large ecosystem of Diffusers and ComfyUI workflows.
- Minimum GPU memory
- 24GB planning floor with an optimised or offload workflow
- Licence
- Stability AI Community Licence
- Useful for
- Image generation
OpenAI
GPT-OSS 20B
A 21-billion-parameter sparse reasoning model with 3.6 billion active parameters, native tool use and an MXFP4 release intended for local or specialised work.
- Minimum GPU memory
- 16GB on one GPU
- Licence
- Apache License 2.0
- Useful for
- Language & reasoning, Coding & agents
Qwen
Qwen3.5 4B
A compact 4-billion-parameter multimodal model with native vision, reasoning, coding and broad multilingual support.
- Minimum GPU memory
- 16GB on one GPU
- Licence
- Apache License 2.0
- Useful for
- Language & reasoning, Coding & agents, Vision & OCR
Qwen
Qwen3.5 27B
A 27-billion-parameter dense multimodal model for reasoning, coding, agents and visual understanding across 201 languages and dialects.
- Minimum GPU memory
- 64GB on one GPU
- Licence
- Apache License 2.0
- Useful for
- Language & reasoning, Coding & agents, Vision & OCR
Google DeepMind
Gemma 4 12B
Google DeepMind's 12-billion-parameter instruction-tuned Gemma 4 model, supporting text, images, video and audio input with text output.
- Minimum GPU memory
- 32GB on one GPU
- Licence
- Apache License 2.0
- Useful for
- Language & reasoning, Vision & OCR, Speech & audio
Google DeepMind
Gemma 4 31B
Google DeepMind's 30.7-billion-parameter instruction-tuned multimodal model for text and image understanding with a 256K context window.
- Minimum GPU memory
- 64GB on one GPU
- Licence
- Apache License 2.0
- Useful for
- Language & reasoning, Vision & OCR
Qwen
Qwen3-TTS 1.7B
A multilingual text-to-speech model with streaming and non-streaming generation, instruction-controlled delivery and nine supplied voice timbres.
- Minimum GPU memory
- 8GB on one GPU
- Licence
- Apache License 2.0
- Useful for
- Speech & audio
Qwen
Qwen3-Reranker 8B
An 8-billion-parameter multilingual reranker for scoring retrieved passages across more than 100 natural and programming languages.
- Minimum GPU memory
- 24GB on one GPU
- Licence
- Apache License 2.0
- Useful for
- Embeddings & RAG
Black Forest Labs
FLUX.2 Klein 9B
A 9-billion-parameter rectified-flow image model for text-to-image generation and multi-reference image editing.
- Minimum GPU memory
- 64GB on one GPU
- Licence
- FLUX Non-Commercial License
- Useful for
- Image generation
StepFun
Step-Audio 2 Mini
An end-to-end audio-language model for spoken conversation, audio understanding, paralinguistic cues and audio tool use.
- Minimum GPU memory
- 24GB on one GPU
- Licence
- Apache License 2.0
- Useful for
- Speech & audio, Language & reasoning
More models through multi-GPU configuration
This system has enough installed GPU memory for these models when supported software divides them across its GPUs. We confirm the runtime, topology, context and performance during sizing.
Compatibility shown here is a hardware-memory screen for the named checkpoint. Before purchase, specify the exact model version, precision, runtime, context, batch, concurrency and required response time.
System details
System specification and service scope
Hardware, software, installation needs and support are brought together around the workload this system needs to handle.
Workloads
Work this system is designed to handle
Performance is tested with representative files, prompts, context, resolution, user count and response-time requirements.
01
Retrieval-augmented generation
Retrieval quality, citation support, permissions, refusal behaviour and latency against a versioned corpus and question set; document count alone is not a hardware metric.
02
Software-engineering assistance
Task correctness, test pass rate, unsafe-change rate, reviewer effort and useful response time on a versioned repository evaluation set.
03
Speech to text
Word error rate, real-time factor and failure rate on a versioned, representative audio set.
04
Embedding
Vectors per second, query latency and retrieval-quality metric on a versioned corpus with the exact embedding checkpoint and dimensions.
05
Vision and OCR
Field accuracy or character error rate plus latency on a versioned, permission-safe image and document set.
06
Image generation
Latency percentiles, images per second, peak memory, stability and accepted-output rate at exact checkpoint, resolution, steps, sampler, CFG and batch.
07
Video generation
Clip latency, clips per hour, peak memory, stability and accepted-output rate at exact checkpoint, dimensions, frames, frame rate, duration, steps and batch.
08
GPU rendering
Frame latency, throughput, errors, wall power and temperature for an immutable scene using the exact renderer, version, resolution and sample count.
09
Fine-tuning
Completed steps, time, peak memory, loss and held-out evaluation result with exact base checkpoint, method, trainable parameters, dataset, sequence length and batch.
10
Language-model inference
Quality pass rate, TTFT, decode and aggregate throughput, latency percentiles, errors, memory, power and temperature at disclosed context, batch and concurrency.
Hardware
Hardware specification
Core hardware facts for the listed configuration. CPU, memory, storage and networking are selected to suit the workload and deployment environment.
| Item | Specification |
|---|---|
| Physical GPU memory total | 768GB across 8 GPUs |
| Per-GPU memory | 96GB |
| GPU interconnect | PCIe; no NVLink claim |
Installation
Power, cooling and placement
The room, power supply, heat rejection, noise tolerance and service access must suit the final system.
Site requirements
- A suitable 19-inch rack, rail depth, handling route and secure operating location
- A qualified electrical design based on the final PSU population and measured load
- Cooling and heat-rejection capacity for sustained accelerator operation
- Appropriate switching, cabling, remote management and network segmentation
- A named operational owner or contracted support route
Delivery and support
Included service and separate responsibilities
GPU Servers supplies a configured, tested and documented system for collection or agreed delivery.
Included
- Documented workload and site-fit review
- Itemised hardware specification before procurement
- Configuration, burn-in and agreed smoke-test evidence
- Asset list, administrator handover notes and user quick-start material
- Collection or the quoted kerbside or pallet-delivery route
- Remote onboarding and 30-day configuration-defect support
Separate service or customer responsibility
- Building electrical work, rack, UPS, cooling or structured cabling
- Nationwide on-site installation unless separately quoted
- Migration of customer data, every integration or every application
- Continuous managed operations, security monitoring or a 24-hour support agreement
- Third-party model, API, marketplace or software charges
- A compliance certificate, performance guarantee or income guarantee
Explore Enterprise 768 hardware and system design.
The system view shows the hardware platform, while the technical diagrams explain memory, data boundaries, queues, power and handover.
Software and security
The usable product is more than the chassis.
The software stack stays deliberately small. The handover identifies versions, licences, access controls and ongoing responsibilities.
Software baseline
- Ubuntu LTS with the installed version stated in the handover
- NVIDIA driver, CUDA components and container support matched to the supplied hardware
- Docker Engine and NVIDIA Container Toolkit
- One primary model server selected from Ollama, vLLM, SGLang, TensorRT-LLM or llama.cpp for the accepted workload
- Open WebUI or another agreed browser interface
- Named authentication, TLS and reverse-proxy approach
- GPU, node and service monitoring with an agreed log-retention period
- Documented software versions, model sources and licence information
Security ownership
- GPU RIGS baseline
- Initial operating-system state, named administrator handover, host firewall baseline, agreed access route, secrets transfer and documented update state.
- Customer or contracted operator
- User lifecycle, network and VPN policy, backups, monitoring review, patch approval, incident response, data governance and lawful use.
- Agreed responsibilities
- Model and software licence checks, retention and logging choices, recovery test, service targets and a named owner for every recurring task.
Local infrastructure can reduce disclosure to external AI APIs. It does not automatically make the service secure, accurate, confidential or UK GDPR compliant.
Testing before handover
The system is tested as a complete machine.
Burn-in covers the complete build. Agreed workload checks use representative inputs and stated conditions.
- 01
Record the itemised hardware specification, serial numbers and firmware versions.
- 02
Run at least 24 hours of GPU, CPU, memory and storage stress testing.
- 03
Capture temperature, fan, error, health and wall-power evidence under the agreed load.
- 04
Test cold boot, restart and the available remote-management route.
- 05
Check drive health, network throughput and the container and GPU runtime.
- 06
Run model and workload smoke tests using the agreed representative inputs.
- 07
Record any measured speed only with the model, quantisation, context, concurrency and runtime disclosed.
- 08
Test an agreed fault or recovery route and include the result in the handover.
How performance figures are presented
Any stated speed, latency, quality, power or capacity result identifies the system, model, workload and test conditions used.
See how systems are testedCommercial reality
Tax and spare capacity are supporting questions.
Neither belongs in a guaranteed saving or payback claim. The purchase must stand on the accepted workload, control case and operating plan.
Finance and capital allowances
- The displayed figure covers the complete GPU Servers private AI deployment described on this page, not unconfigured hardware.
- Displayed figures are guide prices. The final total is confirmed before purchase.
- Qualifying equipment may be plant and machinery for capital-allowance purposes. The buyer's accountant decides eligibility and timing.
- A third-party lease or hire-purchase route may be explored after partner validation and credit approval. GPU RIGS is not presented as a lender.
- There is no generic capital-gains advantage and no automatic research and development relief because the equipment supports AI.
Optional idle capacity
- Marketplace mode is off by default and is excluded from the purchase case.
- A separate environment, no customer data mounts, network controls and a local kill switch would be required.
- The customer, insurer, supplier warranty and marketplace terms must permit the intended use.
- Vast.ai, Render, Golem and direct batch work do not guarantee acceptance, demand, rate or income.
- Any dated estimate must identify its source date; a pilot must report achieved utilisation, fees, electricity, cooling, faults and operator time.
Limits and alternatives
A good specification leaves room for “no”.
The fit check can recommend a smaller system, hosted service, hybrid route, custom build or no purchase. That is preferable to choosing hardware that does not suit the work.
Package boundaries
- Aggregate GPU memory is not one universal memory pool.
- Confirm model speed, usable context, concurrency and multi-GPU efficiency with the exact software and workload.
- Rack, electrical work, UPS, cooling, cabling, colocation and on-site installation require a separate scope.
- 768GB is aggregate capacity across eight 96GB GPUs.
- This is a PCIe architecture, not an HGX system and not one pooled 768GB memory space.
- No Kimi K3 fit, context, throughput, concurrency or fine-tuning result has been measured.
- Aggregate GPU memory is not one universal memory pool.
- Confirm model speed, usable context, concurrency and multi-GPU efficiency with the exact software and workload.
- Rack, electrical work, UPS, cooling, cabling, colocation and on-site installation require a separate scope.
Choose the smaller system when it meets the accepted workload and resilience requirement.
Consider this route Custom specificationUse when storage, network, resilience, colocation or service design differs from the reference route.
Consider this route Hosted or hybrid AIMay be preferable for bursty demand or where the organisation cannot own the facility and operating burden.
Consider this routeQuestions answered
Enterprise 768 questions that affect the order
Answers cover pricing, specifications, warranty, workload fit and the service supplied.
Does Enterprise 768 provide 768GB as one memory pool?
Only a single-GPU system provides its stated GPU memory on one card. Multi-GPU totals are aggregate physical capacity; usable sharding depends on the exact model, runtime and topology.
Which models are supported?
The compatibility section lists models whose published, calculated or estimated minimum memory fits the hardware. Confirm speed, usable context and concurrency with the exact model version and serving software.
What does the displayed price include?
The displayed price covers the complete configured GPU Servers deployment described on the page. Systems marked pricing on request are configured around the selected workload, facility and service needs.
Can spare capacity earn income?
It may be evaluated as an optional isolated secondary use. Marketplace acceptance, utilisation, rates, fees, energy and income can change and are never guaranteed or included in the purchase case.
Prepare the next decision
Specify the work before the parts.
Tell us about the workload, data boundary, users, site and success criteria. No confidential documents or credentials are needed.
Guide prices cover the system and service scope described on this page. The final total is confirmed before purchase.