Best larger-system route
Studio 96
The smallest system in the range that reaches this model's preferred working allowance.
Explore Studio 96
Qwen model
The official BF16 repository is about 71.9GB before runtime and cache overhead. A 96GB single GPU is the cleanest starting point in this range. Smaller 32GB products need a named quantised build, reduced context and separate validation.
Qwen/Qwen3.6-35B-A3B995ad96eacd9Qwen/Qwen3.6-35B-A3B Buyer verdict
Qwen3.6 35B-A3B suits coding assistants, tool use, long technical documents, image-aware applications and mixed English/Chinese work where one high-memory workstation is preferable to a frontier rack.
These figures apply to the named model version. Quantisation, fine-tuning, context length, image resolution, batch size and serving software can materially change the hardware needed.
Hardware requirements
The 71.9GB repository total leaves limited headroom. Context, image tokens and runtime allocations must be bounded.
This is a GPU Servers planning recommendation, not a Qwen benchmark or full-context promise.
Technical specification
Specifications shown for source version
995ad96eacd98c81ed38be0c5b274b04031597b0, updated
24 April 2026.
Product compatibility
A system is listed as fitting when its GPU memory meets the requirement shown above. Speed, usable context, batch size and concurrent users still need testing with the final model and software configuration.
Best larger-system route
The smallest system in the range that reaches this model's preferred working allowance.
Explore Studio 96sme workstation
This system provides 32GB per GPU and 32GB in total. The model requires 80GB on one GPU.
sme workstation
This system provides 32GB per GPU and 64GB in total. The model requires 80GB on one GPU.
sme workstation
Also meets this model page's recommended working allowance.
Available GPU memory exceeds the calculated or estimated requirement. Confirm the final precision, software, workload and performance before purchase.
sme workstation
Also meets this model page's recommended working allowance.
Available GPU memory exceeds the calculated or estimated requirement. Confirm the final precision, software, workload and performance before purchase.
pcie rack
This system provides 32GB per GPU and 128GB in total. The model requires 80GB on one GPU.
pcie rack
This system provides 32GB per GPU and 256GB in total. The model requires 80GB on one GPU.
pcie rack
Also meets this model page's recommended working allowance.
Available GPU memory exceeds the calculated or estimated requirement. Confirm the final precision, software, workload and performance before purchase.
pcie rack
Also meets this model page's recommended working allowance.
Available GPU memory exceeds the calculated or estimated requirement. Confirm the final precision, software, workload and performance before purchase.
pcie rack
Also meets this model page's recommended working allowance.
Available GPU memory exceeds the calculated or estimated requirement. Confirm the final precision, software, workload and performance before purchase.
frontier partner
Also meets this model page's recommended working allowance.
Available GPU memory exceeds the calculated or estimated requirement. Confirm the final precision, software, workload and performance before purchase.
frontier partner
Also meets this model page's recommended working allowance.
Available GPU memory exceeds the calculated or estimated requirement. Confirm the final precision, software, workload and performance before purchase.
Deployment reality
Serving software
Before installation
System requirements
Performance: speed, usable context and concurrency depend on the selected system, software and workload. No benchmark is quoted on this page.
Commercial and legal boundary
Commercial use: permitted
Apache 2.0 permits commercial use, subject to notices and the intended application's legal and governance requirements.
Always retain the applicable notices and recheck the live terms for the intended organisation, territory, use and distribution route. Obtain legal advice where required; the official licence governs use.
Read the official licenceTechnical questions
A smaller quantised conversion may load, but the named BF16 repository is about 71.9GB. A 32GB result must identify the conversion, context and quality trade-off and is not treated as the official-checkpoint fit.
Studio 96 is the first clean single-GPU planning fit. Studio 192 provides a second worker or testing capacity, while rack systems become relevant for multiple endpoints or queues.
No. The model weights, KV cache, visual tokens, batch and runtime share GPU memory. Full-context use needs a separate memory and latency acceptance test.
GPU memory is only one part of the system. Storage, data access, serving software, monitoring and administrator handover also affect a reliable deployment.
Privacy & cookies. Google Analytics measures site use so we can improve it. Read our Cookie Notice and Privacy Notice. To object, disable cookies in your browser.