OpenAI model

GPT-OSS 20B Hardware Requirements & Server Compatibility

GPT-OSS 20B is a current OpenAI release for lower-latency reasoning, tool use and local specialised assistants. The pinned repository contains approximately 41.274GB of model weights. Our 16GB on one GPU figure is a vendor documented memory screen, while 32GB on one GPU is the safer starting allowance for deployment testing.

Model version
openai/gpt-oss-20b
Family and variant
GPT-OSS · 20B MXFP4
Source version
6cee5e81ee83
Source updated
26 August 2025
OpenAI GPT-OSS 20B openai/gpt-oss-20b
Minimum GPU memory
16GB on one GPU
Recommended hardware
32GB on one GPU
Licence
Apache License 2.0
Useful for
Language & reasoning, Coding & agents
Memory compatibility is a sizing guide. Test the exact model version and workload before choosing hardware.

Buyer verdict

Where GPT-OSS 20B is a sensible fit

Shortlist GPT-OSS 20B when lower-latency reasoning, tool use and local specialised assistants is the priority and the exact licence and runtime suit the organisation. Choose hardware from the recommended allowance, then measure quality and performance on representative work before purchase.

These figures apply to the named model version. Quantisation, fine-tuning, context length, image resolution, batch size and serving software can materially change the hardware needed.

Hardware requirements

Minimum
16GB on one GPU
Recommended
32GB on one GPU

This is the developer's published memory target. Usable context and throughput remain workload-dependent.

This working allowance creates room for runtime allocations and representative workload testing; it is not a performance benchmark.

Technical specification

GPT-OSS 20B model and hardware facts

Specifications shown for source version 6cee5e81ee83917806bbde320786a8fb61efebee, updated 26 August 2025.

Architecture
Sparse mixture of experts
Parameters
21B total / 3.6B active
Context
131,072 tokens
Architecture ceiling; serving capacity requires testing
Modalities
Text → text
Repository weights
41.274GB
Pinned first-party repository file total
Native format
MXFP4 MoE weights

Product compatibility

GPT-OSS 20B compatibility across all 11 GPU systems

Systems that fit without model splitting
11
Systems needing multi-GPU validation
0

A system is listed as fitting when its GPU memory meets the requirement shown above. Speed, usable context, batch size and concurrent users still need testing with the final model and software configuration.

Best larger-system route

Team 32

1 GPUs · 32GB

The smallest system in the range that reaches this model's preferred working allowance.

Explore Team 32

sme workstation

Team 32

Recommended memory route
Per GPU
32GB
Total GPU memory
32GB
GPU count
1

Also meets this model page's recommended working allowance.

Available GPU memory exceeds the model developer's published requirement. Confirm speed, context, batch size and concurrent users before purchase.

sme workstation

Company 64

Recommended memory route
Per GPU
32GB
Total GPU memory
64GB
GPU count
2

Also meets this model page's recommended working allowance.

Available GPU memory exceeds the model developer's published requirement. Confirm speed, context, batch size and concurrent users before purchase.

sme workstation

Studio 96

Recommended memory route
Per GPU
96GB
Total GPU memory
96GB
GPU count
1

Also meets this model page's recommended working allowance.

Available GPU memory exceeds the model developer's published requirement. Confirm speed, context, batch size and concurrent users before purchase.

sme workstation

Studio 192

Recommended memory route
Per GPU
96GB
Total GPU memory
192GB
GPU count
2

Also meets this model page's recommended working allowance.

Available GPU memory exceeds the model developer's published requirement. Confirm speed, context, batch size and concurrent users before purchase.

pcie rack

Value Rack 128

Recommended memory route
Per GPU
32GB
Total GPU memory
128GB
GPU count
4

Also meets this model page's recommended working allowance.

Available GPU memory exceeds the model developer's published requirement. Confirm speed, context, batch size and concurrent users before purchase.

pcie rack

Value Rack 256

Recommended memory route
Per GPU
32GB
Total GPU memory
256GB
GPU count
8

Also meets this model page's recommended working allowance.

Available GPU memory exceeds the model developer's published requirement. Confirm speed, context, batch size and concurrent users before purchase.

pcie rack

Enterprise 384

Recommended memory route
Per GPU
96GB
Total GPU memory
384GB
GPU count
4

Also meets this model page's recommended working allowance.

Available GPU memory exceeds the model developer's published requirement. Confirm speed, context, batch size and concurrent users before purchase.

pcie rack

Enterprise 768

Recommended memory route
Per GPU
96GB
Total GPU memory
768GB
GPU count
8

Also meets this model page's recommended working allowance.

Available GPU memory exceeds the model developer's published requirement. Confirm speed, context, batch size and concurrent users before purchase.

pcie rack

H200 1.1TB

Recommended memory route
Per GPU
141GB
Total GPU memory
1,128GB
GPU count
8

Also meets this model page's recommended working allowance.

Available GPU memory exceeds the model developer's published requirement. Confirm speed, context, batch size and concurrent users before purchase.

frontier partner

Frontier Native 2.3TB

Recommended memory route
Per GPU
288GB
Total GPU memory
2,304GB
GPU count
8

Also meets this model page's recommended working allowance.

Available GPU memory exceeds the model developer's published requirement. Confirm speed, context, batch size and concurrent users before purchase.

frontier partner

Frontier Rack 20TB

Recommended memory route
Per GPU
288GB
Total GPU memory
20,736GB
GPU count
72

Also meets this model page's recommended working allowance.

Available GPU memory exceeds the model developer's published requirement. Confirm speed, context, batch size and concurrent users before purchase.

Deployment reality

Strengths, limits and runtime route

What it is good at

  • Officially described as running within 16GB memory.
  • Native configurable reasoning effort and tool use.
  • Apache 2.0 licensing and several documented local runtimes.
  • Smaller active parameter count than the 120B sibling for lower-latency work.

Where to be cautious

  • The Harmony response format must be preserved by the serving stack.
  • The repository contains alternative files, so its total download size is not the same as the published loaded-memory target.
  • Repository weight size does not include every runtime allocation, KV cache, media encoder, batch or concurrent session.
  • No GPU Servers benchmark result is claimed until this exact revision has been run under disclosed test conditions.

Serving software

  • vLLM
  • Transformers
  • Ollama
  • llama.cpp

Before installation

  1. Use source version 6cee5e81ee83917806bbde320786a8fb61efebee rather than an unversioned latest branch.
  2. Start with vLLM only after checking support for the exact architecture and numerical format.
  3. Use 32GB on one GPU as the procurement starting point; the lower figure is a minimum-memory screen, not a service-level promise.
  4. Record precision, context, batch, concurrency, container digest, driver and measured latency in the acceptance report.

System requirements

GPU layout
The minimum is screened against one GPU. Additional GPUs provide replicas or throughput unless the chosen runtime supports model parallelism.
Context and cache
131,072 tokens. Context length, KV cache, batch and concurrent sessions must be tested together; the architecture limit is not a throughput guarantee.
System RAM
Plan system RAM above the 41.274GB repository-weight footprint where model staging or CPU offload is required.
Storage
Reserve at least 104GB for the pinned weights, runtime cache and one rollback copy; production datasets and logs are additional.
Serving software
First-party material names vLLM, Transformers, Ollama, llama.cpp. Pin the serving version and container digest because support for recent architectures can change quickly.
Representative workload
Acceptance-test lower-latency reasoning, tool use and local specialised assistants at the intended quality, context, batch, concurrency and response-time target.

Performance: speed, usable context and concurrency depend on the selected system, software and workload. No benchmark is quoted on this page.

Commercial and legal boundary

Apache License 2.0

Commercial use: permitted

OpenAI releases the model under Apache 2.0, permitting commercial use subject to the licence terms and notices.

Always retain the applicable notices and recheck the live terms for the intended organisation, territory, use and distribution route. Obtain legal advice where required; the official licence governs use.

Read the official licence

Official sources

Technical questions

GPT-OSS 20B deployment FAQ

What GPU memory does GPT-OSS 20B need?

Use 16GB on one GPU as the lower memory screen and 32GB on one GPU as the safer starting allowance. The final requirement changes with runtime, context, batch, concurrency and precision.

Which GPU Servers product should I start with for GPT-OSS 20B?

Use the compatibility table to find systems that meet the recommended allowance. A system that only meets the minimum can load the model in principle but may not meet the required context or response time.

Can GPT-OSS 20B be used commercially?

OpenAI releases the model under Apache 2.0, permitting commercial use subject to the licence terms and notices. The linked official licence is authoritative; legal advice may be appropriate for the intended use and distribution route.

What a complete GPT-OSS 20B deployment needs

GPU memory is only one part of the system. Storage, data access, serving software, monitoring and administrator handover also affect a reliable deployment.

Diagram showing an approved request, a local service, an approved store and a policy-controlled data path
Diagram showing an approved request, a local service, an approved store and a policy-controlled data path
A private deployment starts with the permitted data path, access policy and logging boundary. GPU Servers technical illustration.
A private deployment starts with the permitted data path, access policy and logging boundary. GPU Servers technical illustration.
GPU server remote management dashboard with system status, access logs and sensor monitoring panels
GPU server remote management dashboard with system status, access logs and sensor monitoring panels
Supplier screenshot of the platform management interface. The final management features and access policy depend on the ordered system. OEM supplier reference image.
Supplier screenshot of the platform management interface. The final management features and access policy depend on the ordered system. OEM supplier reference image.
Diagram showing approved documents moving through a searchable index to an answer with a source citation
Diagram showing approved documents moving through a searchable index to an answer with a source citation
A retrieval workflow should connect each useful answer to approved source material and defined refusal behaviour. GPU Servers technical illustration.
A retrieval workflow should connect each useful answer to approved source material and defined refusal behaviour. GPU Servers technical illustration.
Diagram of an evidence pack containing an asset schedule, burn-in record, health readings, workload test and admin guide
Diagram of an evidence pack containing an asset schedule, burn-in record, health readings, workload test and admin guide
A complete handover includes the supplied assets, test results, operating instructions and agreed follow-up work. GPU Servers technical illustration.
A complete handover includes the supplied assets, test results, operating instructions and agreed follow-up work. GPU Servers technical illustration.