Best larger-system route
Team 32
The smallest system in the range that reaches this model's preferred working allowance.
Explore Team 32
OpenAI model
GPT-OSS 20B is a current OpenAI release for lower-latency reasoning, tool use and local specialised assistants. The pinned repository contains approximately 41.274GB of model weights. Our 16GB on one GPU figure is a vendor documented memory screen, while 32GB on one GPU is the safer starting allowance for deployment testing.
openai/gpt-oss-20b6cee5e81ee83openai/gpt-oss-20b Buyer verdict
Shortlist GPT-OSS 20B when lower-latency reasoning, tool use and local specialised assistants is the priority and the exact licence and runtime suit the organisation. Choose hardware from the recommended allowance, then measure quality and performance on representative work before purchase.
These figures apply to the named model version. Quantisation, fine-tuning, context length, image resolution, batch size and serving software can materially change the hardware needed.
Hardware requirements
This is the developer's published memory target. Usable context and throughput remain workload-dependent.
This working allowance creates room for runtime allocations and representative workload testing; it is not a performance benchmark.
Technical specification
Specifications shown for source version
6cee5e81ee83917806bbde320786a8fb61efebee, updated
26 August 2025.
Product compatibility
A system is listed as fitting when its GPU memory meets the requirement shown above. Speed, usable context, batch size and concurrent users still need testing with the final model and software configuration.
Best larger-system route
The smallest system in the range that reaches this model's preferred working allowance.
Explore Team 32sme workstation
Also meets this model page's recommended working allowance.
Available GPU memory exceeds the model developer's published requirement. Confirm speed, context, batch size and concurrent users before purchase.
sme workstation
Also meets this model page's recommended working allowance.
Available GPU memory exceeds the model developer's published requirement. Confirm speed, context, batch size and concurrent users before purchase.
sme workstation
Also meets this model page's recommended working allowance.
Available GPU memory exceeds the model developer's published requirement. Confirm speed, context, batch size and concurrent users before purchase.
sme workstation
Also meets this model page's recommended working allowance.
Available GPU memory exceeds the model developer's published requirement. Confirm speed, context, batch size and concurrent users before purchase.
pcie rack
Also meets this model page's recommended working allowance.
Available GPU memory exceeds the model developer's published requirement. Confirm speed, context, batch size and concurrent users before purchase.
pcie rack
Also meets this model page's recommended working allowance.
Available GPU memory exceeds the model developer's published requirement. Confirm speed, context, batch size and concurrent users before purchase.
pcie rack
Also meets this model page's recommended working allowance.
Available GPU memory exceeds the model developer's published requirement. Confirm speed, context, batch size and concurrent users before purchase.
pcie rack
Also meets this model page's recommended working allowance.
Available GPU memory exceeds the model developer's published requirement. Confirm speed, context, batch size and concurrent users before purchase.
pcie rack
Also meets this model page's recommended working allowance.
Available GPU memory exceeds the model developer's published requirement. Confirm speed, context, batch size and concurrent users before purchase.
frontier partner
Also meets this model page's recommended working allowance.
Available GPU memory exceeds the model developer's published requirement. Confirm speed, context, batch size and concurrent users before purchase.
frontier partner
Also meets this model page's recommended working allowance.
Available GPU memory exceeds the model developer's published requirement. Confirm speed, context, batch size and concurrent users before purchase.
Deployment reality
Serving software
Before installation
System requirements
Performance: speed, usable context and concurrency depend on the selected system, software and workload. No benchmark is quoted on this page.
Commercial and legal boundary
Commercial use: permitted
OpenAI releases the model under Apache 2.0, permitting commercial use subject to the licence terms and notices.
Always retain the applicable notices and recheck the live terms for the intended organisation, territory, use and distribution route. Obtain legal advice where required; the official licence governs use.
Read the official licenceTechnical questions
Use 16GB on one GPU as the lower memory screen and 32GB on one GPU as the safer starting allowance. The final requirement changes with runtime, context, batch, concurrency and precision.
Use the compatibility table to find systems that meet the recommended allowance. A system that only meets the minimum can load the model in principle but may not meet the required context or response time.
OpenAI releases the model under Apache 2.0, permitting commercial use subject to the licence terms and notices. The linked official licence is authoritative; legal advice may be appropriate for the intended use and distribution route.
GPU memory is only one part of the system. Storage, data access, serving software, monitoring and administrator handover also affect a reliable deployment.
Privacy & cookies. Google Analytics measures site use so we can improve it. Read our Cookie Notice and Privacy Notice. To object, disable cookies in your browser.