Best larger-system route
Studio 96
The smallest system in the range that reaches this model's preferred working allowance.
Explore Studio 96
OpenAI model
OpenAI states that GPT-OSS 120B fits on one 80GB GPU. In this range, Studio 96 is the first straightforward single-GPU product; 32GB products are appropriate for GPT-OSS 20B rather than the named 120B checkpoint.
openai/gpt-oss-120bb5c939de8f75openai/gpt-oss-120b Buyer verdict
GPT-OSS 120B is well suited to tool use, structured outputs, reasoning and fine-tuning where Apache licensing and a one-80GB-GPU deployment are important.
These figures apply to the named model version. Quantisation, fine-tuning, context length, image resolution, batch size and serving software can materially change the hardware needed.
Hardware requirements
OpenAI documents this as the serving target for GPT-OSS 120B.
Context and concurrency still determine usable service capacity.
Technical specification
Specifications shown for source version
b5c939de8f754692c1647ca79fbf85e8c1e70f8a, updated
26 August 2025.
Product compatibility
A system is listed as fitting when its GPU memory meets the requirement shown above. Speed, usable context, batch size and concurrent users still need testing with the final model and software configuration.
Best larger-system route
The smallest system in the range that reaches this model's preferred working allowance.
Explore Studio 96sme workstation
This system provides 32GB per GPU and 32GB in total. The model requires 80GB on one GPU.
sme workstation
This system provides 32GB per GPU and 64GB in total. The model requires 80GB on one GPU.
sme workstation
Also meets this model page's recommended working allowance.
Available GPU memory exceeds the model developer's published requirement. Confirm speed, context, batch size and concurrent users before purchase.
sme workstation
Also meets this model page's recommended working allowance.
Available GPU memory exceeds the model developer's published requirement. Confirm speed, context, batch size and concurrent users before purchase.
pcie rack
This system provides 32GB per GPU and 128GB in total. The model requires 80GB on one GPU.
pcie rack
This system provides 32GB per GPU and 256GB in total. The model requires 80GB on one GPU.
pcie rack
Also meets this model page's recommended working allowance.
Available GPU memory exceeds the model developer's published requirement. Confirm speed, context, batch size and concurrent users before purchase.
pcie rack
Also meets this model page's recommended working allowance.
Available GPU memory exceeds the model developer's published requirement. Confirm speed, context, batch size and concurrent users before purchase.
pcie rack
Also meets this model page's recommended working allowance.
Available GPU memory exceeds the model developer's published requirement. Confirm speed, context, batch size and concurrent users before purchase.
frontier partner
Also meets this model page's recommended working allowance.
Available GPU memory exceeds the model developer's published requirement. Confirm speed, context, batch size and concurrent users before purchase.
frontier partner
Also meets this model page's recommended working allowance.
Available GPU memory exceeds the model developer's published requirement. Confirm speed, context, batch size and concurrent users before purchase.
Deployment reality
Serving software
Before installation
System requirements
Performance: speed, usable context and concurrency depend on the selected system, software and workload. No benchmark is quoted on this page.
Commercial and legal boundary
Commercial use: permitted
Apache 2.0 permits commercial use subject to notices and the intended application's obligations.
Always retain the applicable notices and recheck the live terms for the intended organisation, territory, use and distribution route. Obtain legal advice where required; the official licence governs use.
Read the official licenceTechnical questions
Studio 96 is the first product in the range with more than the official 80GB GPU floor.
Team 32 is better matched to GPT-OSS 20B. A third-party 120B conversion is a different model version and should not be described as the official 120B release.
OpenAI documents fine-tuning routes, but training memory differs from inference. The method, optimiser, sequence length and adapter or full-weight scope must be sized separately.
GPU memory is only one part of the system. Storage, data access, serving software, monitoring and administrator handover also affect a reliable deployment.
Privacy & cookies. Google Analytics measures site use so we can improve it. Read our Cookie Notice and Privacy Notice. To object, disable cookies in your browser.