SME workstation · one 32GB GPU

Buy the Team 32 RTX 5090 AI Workstation

A first private AI workstation for one team, without a rack.

A selected-supplier RTX 5090 tower for controlled local chat, RAG, coding, speech, image and batch evaluation when the accepted workload fits one 32GB GPU.

Reference platform image
Black mid-tower workstation selected as the Team 32 reference
Inspect the platform
1 × 32GB GPU memory Single-GPU mid-tower AI workstation reference platform GPU configuration 1 × NVIDIA GeForce RTX 5090 32GB About this image Workstation-family reference image · appearance varies by configuration Memory use 32GB is available on one GPU as one physical memory space.

Key buying facts

Team 32 at a glance

Price
£7,500 guide price · complete configured system
GPU configuration
1 × NVIDIA GeForce RTX 5090 32GB 1 GPU
Per-GPU memory
32GB physical VRAM per GPU
GPU memory
32GB one GPU memory space
Physical class
Workstation Single-GPU mid-tower AI workstation

How the memory works: 32GB is available on one GPU as one physical memory space.

Power planning · Workstation

Specified for the ordered configuration and intended facility.

Representative intended fit

A named workload that fits one 32GB GPU

Buyer fit

Start with the reason to own it.

A supported office-compatible starting point with an evidence-led upgrade route only when measured demand justifies it.

Price

£7,500

guide price · complete configured system

Includes workload sizing, the configured system, AI software stack, security baseline, burn-in, agreed workload testing, documentation, remote onboarding and 30-day configuration-defect support.

Guide prices cover the system and service scope described on this page. The final total is confirmed before purchase.

A credible fit

  • One team beginning a controlled local-AI deployment
  • A named workload that fits one 32GB GPU
  • An office that cannot accommodate a rack server

Choose another route when

  • The accepted model or context needs more than 32GB
  • Several heavy services need guaranteed simultaneous capacity
  • The buyer needs redundant power or enterprise remote management

Model compatibility

Models that fit Team 32's GPU memory

32GB is available on one GPU.

The results below compare that hardware with each model's GPU-memory requirement. They do not predict speed, maximum context, image size or concurrent users.

All 33 models in our current model guide are included below. Another version needs its own exact checkpoint, licence, runtime and memory requirement before it can be matched to hardware.

18 models fit without splitting the model across GPUs

Compare every model

Larger-system opportunities

Choose a larger server for 15 additional models

If these models are part of your plan, the routes below show the smallest larger system that best meets their GPU-memory needs. Compare it with Team 32, then size the final configuration around your workload.

Recommended memory route

Frontier Native 2.3TB

8 GPUs · 2,304GB

The smallest larger system in the range that reaches the preferred working allowance for these models.

DeepSeek

DeepSeek-OCR 2

Recommended memory fit

A 3B-class image-to-text model for document optical character recognition and layout-aware text extraction.

Minimum GPU memory
12GB GPU planning floor
Licence
Apache License 2.0
Useful for
Vision & OCR
Recommended memory fit

The largest Qwen3 embedding checkpoint, designed for multilingual dense retrieval, classification, clustering and text matching.

Minimum GPU memory
20GB GPU planning floor
Licence
Apache License 2.0
Useful for
Embeddings & RAG

BAAI

BGE-M3

Recommended memory fit

A multilingual embedding model that supports dense, sparse and multi-vector retrieval with 1,024 dimensions and inputs up to 8,192 tokens.

Minimum GPU memory
4GB GPU planning floor, or CPU for light use
Licence
MIT License
Useful for
Embeddings & RAG
Recommended memory fit

A compact automatic speech recognition checkpoint released in July 2026 for multilingual transcription and audio understanding.

Minimum GPU memory
8GB GPU planning floor
Licence
Apache License 2.0
Useful for
Speech & audio
Recommended memory fit

OpenAI's 1.55B-parameter multilingual speech-recognition and translation checkpoint, widely supported across transcription runtimes.

Minimum GPU memory
6GB GPU planning floor for an optimised inference runtime
Licence
Apache License 2.0
Useful for
Speech & audio
Recommended memory fit

A 4B-class real-time automatic speech recognition model released by Mistral AI for low-latency streaming transcription.

Minimum GPU memory
24GB GPU planning floor
Licence
Apache License 2.0
Useful for
Speech & audio

Black Forest Labs

FLUX.2 Klein 4B

Recommended memory fit

A compact rectified-flow model for text-to-image, image editing and multi-reference work, released under Apache 2.0.

Minimum GPU memory
About 13GB VRAM
Licence
Apache License 2.0
Useful for
Image generation
Minimum memory fit; more headroom advised

A 5B text-and-image-to-video model that supports 720p generation and an official single-GPU offload route.

Minimum GPU memory
24GB VRAM with documented CPU offload settings
Licence
Apache License 2.0
Useful for
Video generation
Recommended memory fit

An 8.3B text-to-video and image-to-video model with 480p and 720p checkpoints, optional super-resolution and distilled workflows.

Minimum GPU memory
14GB VRAM with model offloading
Licence
Tencent Hunyuan Community Licence
Useful for
Video generation

Tencent

Hunyuan3D 2.1

Minimum memory fit; more headroom advised

Tencent's image-to-3D generation pipeline for shape and texture creation, with open model weights and a model-specific community licence.

Minimum GPU memory
24GB single-GPU planning floor
Licence
Tencent Hunyuan Community Licence
Useful for
3D generation
Recommended memory fit

A 12B multimodal safeguard model for classifying text and image prompts and responses against Meta's hazard taxonomy.

Minimum GPU memory
28GB GPU planning floor
Licence
Llama 4 Community Licence
Useful for
Safety & moderation, Vision & OCR
Recommended memory fit

An 8B-class text-to-image model in the Stable Diffusion 3.5 family, with a large ecosystem of Diffusers and ComfyUI workflows.

Minimum GPU memory
24GB planning floor with an optimised or offload workflow
Licence
Stability AI Community Licence
Useful for
Image generation

OpenAI

GPT-OSS 20B

Recommended memory fit

A 21-billion-parameter sparse reasoning model with 3.6 billion active parameters, native tool use and an MXFP4 release intended for local or specialised work.

Minimum GPU memory
16GB on one GPU
Licence
Apache License 2.0
Useful for
Language & reasoning, Coding & agents

Qwen

Qwen3.5 4B

Recommended memory fit

A compact 4-billion-parameter multimodal model with native vision, reasoning, coding and broad multilingual support.

Minimum GPU memory
16GB on one GPU
Licence
Apache License 2.0
Useful for
Language & reasoning, Coding & agents, Vision & OCR

Google DeepMind

Gemma 4 12B

Minimum memory fit; more headroom advised

Google DeepMind's 12-billion-parameter instruction-tuned Gemma 4 model, supporting text, images, video and audio input with text output.

Minimum GPU memory
32GB on one GPU
Licence
Apache License 2.0
Useful for
Language & reasoning, Vision & OCR, Speech & audio
Recommended memory fit

A multilingual text-to-speech model with streaming and non-streaming generation, instruction-controlled delivery and nine supplied voice timbres.

Minimum GPU memory
8GB on one GPU
Licence
Apache License 2.0
Useful for
Speech & audio
Recommended memory fit

An 8-billion-parameter multilingual reranker for scoring retrieved passages across more than 100 natural and programming languages.

Minimum GPU memory
24GB on one GPU
Licence
Apache License 2.0
Useful for
Embeddings & RAG
Recommended memory fit

An end-to-end audio-language model for spoken conversation, audio understanding, paralinguistic cues and audio tool use.

Minimum GPU memory
24GB on one GPU
Licence
Apache License 2.0
Useful for
Speech & audio, Language & reasoning

Compatibility shown here is a hardware-memory screen for the named checkpoint. Before purchase, specify the exact model version, precision, runtime, context, batch, concurrency and required response time.

System details

System specification and service scope

Hardware, software, installation needs and support are brought together around the workload this system needs to handle.

Workloads

Work this system is designed to handle

Performance is tested with representative files, prompts, context, resolution, user count and response-time requirements.

01

Retrieval-augmented generation

Retrieval quality, citation support, permissions, refusal behaviour and latency against a versioned corpus and question set; document count alone is not a hardware metric.

02

Software-engineering assistance

Task correctness, test pass rate, unsafe-change rate, reviewer effort and useful response time on a versioned repository evaluation set.

03

Speech to text

Word error rate, real-time factor and failure rate on a versioned, representative audio set.

04

Embedding

Vectors per second, query latency and retrieval-quality metric on a versioned corpus with the exact embedding checkpoint and dimensions.

05

Image generation

Latency percentiles, images per second, peak memory, stability and accepted-output rate at exact checkpoint, resolution, steps, sampler, CFG and batch.

06

GPU rendering

Frame latency, throughput, errors, wall power and temperature for an immutable scene using the exact renderer, version, resolution and sample count.

Hardware

Hardware specification

Core hardware facts for the listed configuration. CPU, memory, storage and networking are selected to suit the workload and deployment environment.

Team 32 specifications
Item Specification
Platform class Single-GPU mid-tower AI workstation
Chassis class Mid-tower workstation
GPU route 1 × NVIDIA GeForce RTX 5090 32GB
GPU memory 32GB across 1 GPUs
Per-GPU memory 32GB
CPU AMD Ryzen 9 9950X
System memory 64GB DDR5
Primary storage 2TB NVMe SSD

Installation

Power, cooling and placement

The room, power supply, heat rejection, noise tolerance and service access must suit the final system.

Site requirements

  • A ventilated floor or desk-side position with clear intake and exhaust paths
  • A suitable UK circuit and an agreed UPS decision after measured-load review
  • An acoustic and heat check under the accepted workload
  • A named owner for accounts, updates, backups and incident response

Delivery and support

Included service and separate responsibilities

GPU Servers supplies a configured, tested and documented system for collection or agreed delivery.

Included

  • Documented workload and site-fit review
  • Itemised hardware specification before procurement
  • Configuration, burn-in and agreed smoke-test evidence
  • Asset list, administrator handover notes and user quick-start material
  • Collection or the quoted kerbside or pallet-delivery route
  • Remote onboarding and 30-day configuration-defect support

Separate service or customer responsibility

  • Building electrical work, rack, UPS, cooling or structured cabling
  • Nationwide on-site installation unless separately quoted
  • Migration of customer data, every integration or every application
  • Continuous managed operations, security monitoring or a 24-hour support agreement
  • Third-party model, API, marketplace or software charges
  • A compliance certificate, performance guarantee or income guarantee

Explore Team 32 hardware and system design.

The system view shows the hardware platform, while the technical diagrams explain memory, data boundaries, queues, power and handover.

Black mid-tower workstation with a ventilated front panel
Black mid-tower workstation with a ventilated front panel
Supplier reference view of the workstation family. Ordered parts and appearance vary by configuration. UK workstation supplier reference image.
Supplier reference view of the workstation family. Ordered parts and appearance vary by configuration. UK workstation supplier reference image.
Diagram combining model weights, context, cache and active requests into a memory headroom check
Diagram combining model weights, context, cache and active requests into a memory headroom check
Memory fit depends on the workload, context, cache and simultaneous demand, then needs a representative test. GPU Servers technical illustration.
Memory fit depends on the workload, context, cache and simultaneous demand, then needs a representative test. GPU Servers technical illustration.
Diagram showing an approved request, a local service, an approved store and a policy-controlled data path
Diagram showing an approved request, a local service, an approved store and a policy-controlled data path
A private deployment starts with the permitted data path, access policy and logging boundary. GPU Servers technical illustration.
A private deployment starts with the permitted data path, access policy and logging boundary. GPU Servers technical illustration.
Diagram of an evidence pack containing an asset schedule, burn-in record, health readings, workload test and admin guide
Diagram of an evidence pack containing an asset schedule, burn-in record, health readings, workload test and admin guide
A complete handover includes the supplied assets, test results, operating instructions and agreed follow-up work. GPU Servers technical illustration.
A complete handover includes the supplied assets, test results, operating instructions and agreed follow-up work. GPU Servers technical illustration.

Software and security

The usable product is more than the chassis.

The software stack stays deliberately small. The handover identifies versions, licences, access controls and ongoing responsibilities.

Software baseline

  1. Ubuntu LTS with the installed version stated in the handover
  2. NVIDIA driver, CUDA components and container support matched to the supplied hardware
  3. Docker Engine and NVIDIA Container Toolkit
  4. One primary model server selected from Ollama, vLLM, SGLang, TensorRT-LLM or llama.cpp for the accepted workload
  5. Open WebUI or another agreed browser interface
  6. Named authentication, TLS and reverse-proxy approach
  7. GPU, node and service monitoring with an agreed log-retention period
  8. Documented software versions, model sources and licence information

Security ownership

GPU RIGS baseline
Initial operating-system state, named administrator handover, host firewall baseline, agreed access route, secrets transfer and documented update state.
Customer or contracted operator
User lifecycle, network and VPN policy, backups, monitoring review, patch approval, incident response, data governance and lawful use.
Agreed responsibilities
Model and software licence checks, retention and logging choices, recovery test, service targets and a named owner for every recurring task.

Local infrastructure can reduce disclosure to external AI APIs. It does not automatically make the service secure, accurate, confidential or UK GDPR compliant.

Testing before handover

The system is tested as a complete machine.

Burn-in covers the complete build. Agreed workload checks use representative inputs and stated conditions.

  1. 01

    Record the itemised hardware specification, serial numbers and firmware versions.

  2. 02

    Run at least 24 hours of GPU, CPU, memory and storage stress testing.

  3. 03

    Capture temperature, fan, error, health and wall-power evidence under the agreed load.

  4. 04

    Test cold boot, restart and the available remote-management route.

  5. 05

    Check drive health, network throughput and the container and GPU runtime.

  6. 06

    Run model and workload smoke tests using the agreed representative inputs.

  7. 07

    Record any measured speed only with the model, quantisation, context, concurrency and runtime disclosed.

  8. 08

    Test an agreed fault or recovery route and include the result in the handover.

How performance figures are presented

Any stated speed, latency, quality, power or capacity result identifies the system, model, workload and test conditions used.

See how systems are tested

Commercial reality

Tax and spare capacity are supporting questions.

Neither belongs in a guaranteed saving or payback claim. The purchase must stand on the accepted workload, control case and operating plan.

Finance and capital allowances

  • The displayed figure covers the complete GPU Servers private AI deployment described on this page, not unconfigured hardware.
  • Displayed figures are guide prices. The final total is confirmed before purchase.
  • Qualifying equipment may be plant and machinery for capital-allowance purposes. The buyer's accountant decides eligibility and timing.
  • A third-party lease or hire-purchase route may be explored after partner validation and credit approval. GPU RIGS is not presented as a lender.
  • There is no generic capital-gains advantage and no automatic research and development relief because the equipment supports AI.
Read the guarded UK buyer notes

Optional idle capacity

  • Marketplace mode is off by default and is excluded from the purchase case.
  • A separate environment, no customer data mounts, network controls and a local kill switch would be required.
  • The customer, insurer, supplier warranty and marketplace terms must permit the intended use.
  • Vast.ai, Render, Golem and direct batch work do not guarantee acceptance, demand, rate or income.
  • Any dated estimate must identify its source date; a pilot must report achieved utilisation, fees, electricity, cooling, faults and operator time.
Read the spare-capacity guide

Limits and alternatives

A good specification leaves room for “no”.

The fit check can recommend a smaller system, hosted service, hybrid route, custom build or no purchase. That is preferable to choosing hardware that does not suit the work.

Package boundaries

  • A workstation is not presented as a redundant datacentre appliance.
  • Confirm model speed, usable context, concurrency and output quality with the exact software and workload.
  • The selected components and delivery route are matched to the configured system.
  • 32GB is one GPU's physical memory; context, cache, batching and concurrency reduce the usable model envelope.
  • No tokens-per-second, user-count, latency, image or video performance has been measured for this revision.
  • A workstation is not presented as a redundant datacentre appliance.
  • Confirm model speed, usable context, concurrency and output quality with the exact software and workload.
  • The selected components and delivery route are matched to the configured system.

Questions answered

Team 32 questions that affect the order

Answers cover pricing, specifications, warranty, workload fit and the service supplied.

Does Team 32 provide 32GB as one memory pool?

Only a single-GPU system provides its stated GPU memory on one card. Multi-GPU totals are aggregate physical capacity; usable sharding depends on the exact model, runtime and topology.

Which models are supported?

The compatibility section lists models whose published, calculated or estimated minimum memory fits the hardware. Confirm speed, usable context and concurrency with the exact model version and serving software.

What does the displayed price include?

The displayed price covers the complete configured GPU Servers deployment described on the page. Systems marked pricing on request are configured around the selected workload, facility and service needs.

Can spare capacity earn income?

It may be evaluated as an optional isolated secondary use. Marketplace acceptance, utilisation, rates, fees, energy and income can change and are never guaranteed or included in the purchase case.

Prepare the next decision

Specify the work before the parts.

Tell us about the workload, data boundary, users, site and success criteria. No confidential documents or credentials are needed.

Guide prices cover the system and service scope described on this page. The final total is confirmed before purchase.

Decision check

RTX 5090 AI Workstation: Fit, Evidence & Next Steps

Choose a RTX 5090 AI workstation only after checking the workload, GPU memory, topology, power, cooling and service requirements against the proposed build.

Relevant supporting considerations include AI workstation, buy AI workstation, AI workstation price, local AI workstation and 32GB GPU workstation. The final RTX 5090 AI workstation quotation should record the bill of materials, compatibility evidence, acceptance tests, warranty and delivery boundary.