Best larger-system route
Frontier Native 2.3TB
The smallest system in the range that reaches this model's preferred working allowance.
Explore Frontier Native 2.3TB
Mistral AI model
The official release contains about 241.9GB of weights. Treat 256GB aggregate as a tight sharded floor and 384GB as the first practical product envelope in this range.
mistralai/Mistral-Small-4-119B-2603a11f36bebf70mistralai/Mistral-Small-4-119B-2603 Buyer verdict
Mistral Small 4 suits organisations that want one Apache-licensed model for general chat, visual inputs, reasoning and coding-agent work on an enterprise multi-GPU server.
These figures apply to the named model version. Quantisation, fine-tuning, context length, image resolution, batch size and serving software can materially change the hardware needed.
Hardware requirements
This is close to the repository total and leaves little practical context or runtime headroom.
The recommendation favours operating headroom and fewer model-sharding boundaries.
Technical specification
Specifications shown for source version
a11f36bebf709121056b1dbcc943d1c6afbe494d, updated
15 July 2026.
Product compatibility
A system is listed as fitting when its GPU memory meets the requirement shown above. Speed, usable context, batch size and concurrent users still need testing with the final model and software configuration.
Best larger-system route
The smallest system in the range that reaches this model's preferred working allowance.
Explore Frontier Native 2.3TBsme workstation
This system provides 32GB per GPU and 32GB in total. The model requires 256GB aggregate across at least 4 GPUs.
sme workstation
This system provides 32GB per GPU and 64GB in total. The model requires 256GB aggregate across at least 4 GPUs.
sme workstation
This system provides 96GB per GPU and 96GB in total. The model requires 256GB aggregate across at least 4 GPUs.
sme workstation
This system provides 96GB per GPU and 192GB in total. The model requires 256GB aggregate across at least 4 GPUs.
pcie rack
This system provides 32GB per GPU and 128GB in total. The model requires 256GB aggregate across at least 4 GPUs.
pcie rack
Total installed memory is sufficient, but the model must be divided across GPUs. Confirm that the serving software supports this GPU layout and test the required context and speed.
pcie rack
Total installed memory is sufficient, but the model must be divided across GPUs. Confirm that the serving software supports this GPU layout and test the required context and speed.
pcie rack
Total installed memory is sufficient, but the model must be divided across GPUs. Confirm that the serving software supports this GPU layout and test the required context and speed.
pcie rack
Total installed memory is sufficient, but the model must be divided across GPUs. Confirm that the serving software supports this GPU layout and test the required context and speed.
frontier partner
Also meets this model page's recommended working allowance.
Available GPU memory exceeds the calculated or estimated requirement. Confirm the final precision, software, workload and performance before purchase.
frontier partner
Also meets this model page's recommended working allowance.
Available GPU memory exceeds the calculated or estimated requirement. Confirm the final precision, software, workload and performance before purchase.
Deployment reality
Serving software
Before installation
System requirements
Performance: speed, usable context and concurrency depend on the selected system, software and workload. No benchmark is quoted on this page.
Commercial and legal boundary
Commercial use: permitted
Apache 2.0 permits commercial use subject to its notice and attribution provisions.
Always retain the applicable notices and recheck the live terms for the intended organisation, territory, use and distribution route. Obtain legal advice where required; the official licence governs use.
Read the official licenceTechnical questions
Value Rack 256 clears the raw aggregate file-size floor, but its eight separate 32GB cards require supported multi-GPU loading. Enterprise 384 is the first more credible planning fit.
No. Active parameters affect compute per token, but the 119B checkpoint weights still need to be stored and addressed.
Yes, the official model is multimodal. Include image count, resolution and context in the workload test.
GPU memory is only one part of the system. Storage, data access, serving software, monitoring and administrator handover also affect a reliable deployment.
Privacy & cookies. Google Analytics measures site use so we can improve it. Read our Cookie Notice and Privacy Notice. To object, disable cookies in your browser.