Meta model

Llama 4 Scout Hardware Requirements & Compatibility

The official BF16 checkpoint contains about 217.3GB of weights. A sharded 256GB server is the tight arithmetic floor, while 384GB provides a more credible operating envelope.

Model version
meta-llama/Llama-4-Scout-17B-16E-Instruct
Family and variant
Llama 4 Scout · Llama-4-Scout-17B-16E-Instruct
Source version
92f3b1597a19
Source updated
22 May 2025
Meta Llama 4 Scout meta-llama/Llama-4-Scout-17B-16E-Instruct
Minimum GPU memory
At least 240GB aggregate for BF16 weights and minimal overhead
Recommended hardware
384GB aggregate for a less constrained sharded deployment
Licence
Llama 4 Community Licence
Useful for
Language & reasoning, Vision & OCR
Memory compatibility is a sizing guide. Test the exact model version and workload before choosing hardware.

Buyer verdict

Where Llama 4 Scout is a sensible fit

Llama 4 Scout is relevant where the team wants the Llama ecosystem, multimodal input and long-context experiments, and can accept Meta's gated licence and multi-GPU deployment needs.

These figures apply to the named model version. Quantisation, fine-tuning, context length, image resolution, batch size and serving software can materially change the hardware needed.

Hardware requirements

Minimum
At least 240GB aggregate for BF16 weights and minimal overhead
Recommended
384GB aggregate for a less constrained sharded deployment

A quantised build can change this floor, but it is a separate artefact.

Long contexts and images can require substantially more memory.

Technical specification

Llama 4 Scout model and hardware facts

Specifications shown for source version 92f3b1597a195b523d8d9e5700e57e4fbb8f20d3, updated 22 May 2025.

Architecture
MoE, 109B total / 17B active
Modalities
Text + image → text
Repository weights
217.28GB BF16
Access
Gated model
Licence
Llama 4 Community Licence
Serving
Runtime-specific multimodal support required

Product compatibility

Llama 4 Scout compatibility across all 11 GPU systems

Systems that fit without model splitting
2
Systems needing multi-GPU validation
4

A system is listed as fitting when its GPU memory meets the requirement shown above. Speed, usable context, batch size and concurrent users still need testing with the final model and software configuration.

sme workstation

Team 32

Larger-memory system recommended
Per GPU
32GB
Total GPU memory
32GB
GPU count
1

This system provides 32GB per GPU and 32GB in total. The model requires 240GB aggregate across at least 4 GPUs.

For this model, explore Frontier Native 2.3TB .

sme workstation

Company 64

Larger-memory system recommended
Per GPU
32GB
Total GPU memory
64GB
GPU count
2

This system provides 32GB per GPU and 64GB in total. The model requires 240GB aggregate across at least 4 GPUs.

For this model, explore Frontier Native 2.3TB .

sme workstation

Studio 96

Larger-memory system recommended
Per GPU
96GB
Total GPU memory
96GB
GPU count
1

This system provides 96GB per GPU and 96GB in total. The model requires 240GB aggregate across at least 4 GPUs.

For this model, explore Frontier Native 2.3TB .

sme workstation

Studio 192

Larger-memory system recommended
Per GPU
96GB
Total GPU memory
192GB
GPU count
2

This system provides 96GB per GPU and 192GB in total. The model requires 240GB aggregate across at least 4 GPUs.

For this model, explore Frontier Native 2.3TB .

pcie rack

Value Rack 128

Larger-memory system recommended
Per GPU
32GB
Total GPU memory
128GB
GPU count
4

This system provides 32GB per GPU and 128GB in total. The model requires 240GB aggregate across at least 4 GPUs.

For this model, explore Frontier Native 2.3TB .

pcie rack

Value Rack 256

Multi-GPU route available
Per GPU
32GB
Total GPU memory
256GB
GPU count
8

Total installed memory is sufficient, but the model must be divided across GPUs. Confirm that the serving software supports this GPU layout and test the required context and speed.

pcie rack

Enterprise 384

Multi-GPU route available
Per GPU
96GB
Total GPU memory
384GB
GPU count
4

Total installed memory is sufficient, but the model must be divided across GPUs. Confirm that the serving software supports this GPU layout and test the required context and speed.

pcie rack

Enterprise 768

Multi-GPU route available
Per GPU
96GB
Total GPU memory
768GB
GPU count
8

Total installed memory is sufficient, but the model must be divided across GPUs. Confirm that the serving software supports this GPU layout and test the required context and speed.

pcie rack

H200 1.1TB

Multi-GPU route available
Per GPU
141GB
Total GPU memory
1,128GB
GPU count
8

Total installed memory is sufficient, but the model must be divided across GPUs. Confirm that the serving software supports this GPU layout and test the required context and speed.

frontier partner

Frontier Native 2.3TB

Recommended memory route
Per GPU
288GB
Total GPU memory
2,304GB
GPU count
8

Also meets this model page's recommended working allowance.

Available GPU memory exceeds the calculated or estimated requirement. Confirm the final precision, software, workload and performance before purchase.

frontier partner

Frontier Rack 20TB

Recommended memory route
Per GPU
288GB
Total GPU memory
20,736GB
GPU count
72

Also meets this model page's recommended working allowance.

Available GPU memory exceeds the calculated or estimated requirement. Confirm the final precision, software, workload and performance before purchase.

Deployment reality

Strengths, limits and runtime route

What it is good at

  • Large Llama ecosystem and broad serving-framework interest.
  • Text and image input with a sparse 16-expert architecture.
  • 17B active parameters despite a 109B total checkpoint.
  • Suitable for multimodal retrieval and long-document evaluation when the context is carefully bounded.

Where to be cautious

  • Access is gated and governed by the Llama 4 community licence.
  • The official BF16 weights are substantially larger than one 96GB GPU.
  • Advertised context ceilings do not establish practical cache, latency or quality on a given server.
  • Commercial terms, acceptable-use policy and attribution need review before deployment.

Serving software

  • vLLM after version check
  • Transformers
  • SGLang after version check

Before installation

  1. Accept Meta's terms through the intended legal entity before downloading the checkpoint.
  2. Use an exact runtime version with documented Llama 4 multimodal support.
  3. Begin with one image and bounded context in acceptance tests.
  4. Do not substitute a community quantisation without recording its publisher, revision and licence inheritance.

System requirements

GPU layout
Plan for at least 4 GPUs and validate model parallelism on the final topology.
Context and cache
The model context ceiling is not a guaranteed serving target. KV cache, batch size and concurrent sessions need separate capacity tests.
System RAM
Size system memory for model loading, runtime overhead, preprocessing and any CPU offload used by the final configuration.
Storage
Allow space for the pinned checkpoint, runtime images, caches, logs and at least one rollback version.
Serving software
Validate the exact checkpoint with vLLM after version check, Transformers, SGLang after version check before acceptance.
Representative workload
Benchmark representative prompts or media at the required context, quality, latency and concurrency.

Performance: speed, usable context and concurrency depend on the selected system, software and workload. No benchmark is quoted on this page.

Commercial and legal boundary

Llama 4 Community Licence

Commercial use: conditional

Commercial use is subject to Meta's model-specific terms, acceptable-use policy and access approval.

Always retain the applicable notices and recheck the live terms for the intended organisation, territory, use and distribution route. Obtain legal advice where required; the official licence governs use.

Read the official licence

Official sources

Technical questions

Llama 4 Scout deployment FAQ

Can Llama 4 Scout run on 96GB?

Not as the named BF16 checkpoint. Community quantisations may reduce memory, but the exact conversion and runtime need a separate record.

Is Value Rack 256 compatible?

Its aggregate total clears a tight arithmetic floor, but the memory is split across eight PCIe GPUs. Treat it as conditional until the exact runtime and topology pass.

Is the Llama licence open source?

It is an open-weight community licence with its own conditions, not Apache 2.0 or MIT. Review the current terms for the intended use.

What a complete Llama 4 Scout deployment needs

GPU memory is only one part of the system. Storage, data access, serving software, monitoring and administrator handover also affect a reliable deployment.

Diagram showing an approved request, a local service, an approved store and a policy-controlled data path
Diagram showing an approved request, a local service, an approved store and a policy-controlled data path
A private deployment starts with the permitted data path, access policy and logging boundary. GPU Servers technical illustration.
A private deployment starts with the permitted data path, access policy and logging boundary. GPU Servers technical illustration.
GPU server remote management dashboard with system status, access logs and sensor monitoring panels
GPU server remote management dashboard with system status, access logs and sensor monitoring panels
Supplier screenshot of the platform management interface. The final management features and access policy depend on the ordered system. OEM supplier reference image.
Supplier screenshot of the platform management interface. The final management features and access policy depend on the ordered system. OEM supplier reference image.
Diagram showing approved documents moving through a searchable index to an answer with a source citation
Diagram showing approved documents moving through a searchable index to an answer with a source citation
A retrieval workflow should connect each useful answer to approved source material and defined refusal behaviour. GPU Servers technical illustration.
A retrieval workflow should connect each useful answer to approved source material and defined refusal behaviour. GPU Servers technical illustration.
Diagram of an evidence pack containing an asset schedule, burn-in record, health readings, workload test and admin guide
Diagram of an evidence pack containing an asset schedule, burn-in record, health readings, workload test and admin guide
A complete handover includes the supplied assets, test results, operating instructions and agreed follow-up work. GPU Servers technical illustration.
A complete handover includes the supplied assets, test results, operating instructions and agreed follow-up work. GPU Servers technical illustration.