Z.ai model

GLM-5 Hardware Requirements & Server Compatibility

GLM-5 is a current Z.ai release for frontier systems engineering, long-horizon agents and complex reasoning. The pinned repository contains approximately 1,507.736GB of model weights. Our 1536GB aggregate across at least 8 GPUs figure is a calculated memory screen, while 2304GB aggregate across at least 8 GPUs is the safer starting allowance for deployment testing.

Model version
zai-org/GLM-5
Family and variant
GLM · 5
Source version
4e6698ba8e85
Source updated
5 April 2026
Z.ai GLM-5 zai-org/GLM-5
Minimum GPU memory
1536GB aggregate across at least 8 GPUs
Recommended hardware
2304GB aggregate across at least 8 GPUs
Licence
MIT License
Useful for
Language & reasoning, Coding & agents
Memory compatibility is a sizing guide. Test the exact model version and workload before choosing hardware.

Buyer verdict

Where GLM-5 is a sensible fit

Shortlist GLM-5 when frontier systems engineering, long-horizon agents and complex reasoning is the priority and the exact licence and runtime suit the organisation. Choose hardware from the recommended allowance, then measure quality and performance on representative work before purchase.

These figures apply to the named model version. Quantisation, fine-tuning, context length, image resolution, batch size and serving software can materially change the hardware needed.

Hardware requirements

Minimum
1536GB aggregate across at least 8 GPUs
Recommended
2304GB aggregate across at least 8 GPUs

This allowance is derived from the pinned artifact and leaves only limited runtime headroom.

This working allowance creates room for runtime allocations and representative workload testing; it is not a performance benchmark.

Technical specification

GLM-5 model and hardware facts

Specifications shown for source version 4e6698ba8e85059d749020e3c4d2123719f23926, updated 5 April 2026.

Architecture
Sparse MoE with DSA attention
Parameters
744B total / 40B active
Context
202,752 tokens
Architecture ceiling; serving capacity requires testing
Modalities
Text → text
Repository weights
1,507.736GB
Pinned first-party repository file total
Native format
BF16 safetensors

Product compatibility

GLM-5 compatibility across all 11 GPU systems

Systems that fit without model splitting
0
Systems needing multi-GPU validation
2

A system is listed as fitting when its GPU memory meets the requirement shown above. Speed, usable context, batch size and concurrent users still need testing with the final model and software configuration.

First multi-GPU opportunity

Frontier Native 2.3TB

8 GPUs · 2,304GB

The first system with enough installed GPU memory when supported software divides the model across its GPUs. We confirm software support, GPU layout, context and performance during sizing.

Explore Frontier Native 2.3TB

sme workstation

Team 32

Larger-memory system recommended
Per GPU
32GB
Total GPU memory
32GB
GPU count
1

This system provides 32GB per GPU and 32GB in total. The model requires 1536GB aggregate across at least 8 GPUs.

For this model, explore Frontier Native 2.3TB .

sme workstation

Company 64

Larger-memory system recommended
Per GPU
32GB
Total GPU memory
64GB
GPU count
2

This system provides 32GB per GPU and 64GB in total. The model requires 1536GB aggregate across at least 8 GPUs.

For this model, explore Frontier Native 2.3TB .

sme workstation

Studio 96

Larger-memory system recommended
Per GPU
96GB
Total GPU memory
96GB
GPU count
1

This system provides 96GB per GPU and 96GB in total. The model requires 1536GB aggregate across at least 8 GPUs.

For this model, explore Frontier Native 2.3TB .

sme workstation

Studio 192

Larger-memory system recommended
Per GPU
96GB
Total GPU memory
192GB
GPU count
2

This system provides 96GB per GPU and 192GB in total. The model requires 1536GB aggregate across at least 8 GPUs.

For this model, explore Frontier Native 2.3TB .

pcie rack

Value Rack 128

Larger-memory system recommended
Per GPU
32GB
Total GPU memory
128GB
GPU count
4

This system provides 32GB per GPU and 128GB in total. The model requires 1536GB aggregate across at least 8 GPUs.

For this model, explore Frontier Native 2.3TB .

pcie rack

Value Rack 256

Larger-memory system recommended
Per GPU
32GB
Total GPU memory
256GB
GPU count
8

This system provides 32GB per GPU and 256GB in total. The model requires 1536GB aggregate across at least 8 GPUs.

For this model, explore Frontier Native 2.3TB .

pcie rack

Enterprise 384

Larger-memory system recommended
Per GPU
96GB
Total GPU memory
384GB
GPU count
4

This system provides 96GB per GPU and 384GB in total. The model requires 1536GB aggregate across at least 8 GPUs.

For this model, explore Frontier Native 2.3TB .

pcie rack

Enterprise 768

Larger-memory system recommended
Per GPU
96GB
Total GPU memory
768GB
GPU count
8

This system provides 96GB per GPU and 768GB in total. The model requires 1536GB aggregate across at least 8 GPUs.

For this model, explore Frontier Native 2.3TB .

pcie rack

H200 1.1TB

Larger-memory system recommended
Per GPU
141GB
Total GPU memory
1,128GB
GPU count
8

This system provides 141GB per GPU and 1128GB in total. The model requires 1536GB aggregate across at least 8 GPUs.

For this model, explore Frontier Native 2.3TB .

frontier partner

Frontier Native 2.3TB

Multi-GPU route available
Per GPU
288GB
Total GPU memory
2,304GB
GPU count
8

Total installed memory is sufficient, but the model must be divided across GPUs. Confirm that the serving software supports this GPU layout and test the required context and speed.

frontier partner

Frontier Rack 20TB

Multi-GPU route available
Per GPU
288GB
Total GPU memory
20,736GB
GPU count
72

Total installed memory is sufficient, but the model must be divided across GPUs. Confirm that the serving software supports this GPU layout and test the required context and speed.

Deployment reality

Strengths, limits and runtime route

What it is good at

  • Designed for long-horizon agentic engineering work.
  • 40B active parameters from a 744B sparse model.
  • Approximately 200K context.
  • MIT licensing.

Where to be cautious

  • The BF16 release is a frontier multi-GPU checkpoint.
  • Only the two largest current products clear the minimum aggregate-memory screen.
  • Repository weight size does not include every runtime allocation, KV cache, media encoder, batch or concurrent session.
  • No GPU Servers benchmark result is claimed until this exact revision has been run under disclosed test conditions.

Serving software

  • vLLM
  • SGLang
  • Transformers

Before installation

  1. Use source version 4e6698ba8e85059d749020e3c4d2123719f23926 rather than an unversioned latest branch.
  2. Start with vLLM only after checking support for the exact architecture and numerical format.
  3. Use 2304GB aggregate across at least 8 GPUs as the procurement starting point; the lower figure is a minimum-memory screen, not a service-level promise.
  4. Record precision, context, batch, concurrency, container digest, driver and measured latency in the acceptance report.

System requirements

GPU layout
The 1536GB aggregate across at least 8 GPUs screen assumes supported tensor or pipeline parallel loading. Aggregate memory is not automatically pooled.
Context and cache
202,752 tokens. Context length, KV cache, batch and concurrent sessions must be tested together; the architecture limit is not a throughput guarantee.
System RAM
Plan system RAM above the 1,507.736GB repository-weight footprint where model staging or CPU offload is required.
Storage
Reserve at least 3770GB for the pinned weights, runtime cache and one rollback copy; production datasets and logs are additional.
Serving software
First-party material names vLLM, SGLang, Transformers. Pin the serving version and container digest because support for recent architectures can change quickly.
Representative workload
Acceptance-test frontier systems engineering, long-horizon agents and complex reasoning at the intended quality, context, batch, concurrency and response-time target.

Performance: speed, usable context and concurrency depend on the selected system, software and workload. No benchmark is quoted on this page.

Commercial and legal boundary

MIT License

Commercial use: permitted

The official GLM-5 checkpoint is MIT licensed.

Always retain the applicable notices and recheck the live terms for the intended organisation, territory, use and distribution route. Obtain legal advice where required; the official licence governs use.

Read the official licence

Official sources

Technical questions

GLM-5 deployment FAQ

What GPU memory does GLM-5 need?

Use 1536GB aggregate across at least 8 GPUs as the lower memory screen and 2304GB aggregate across at least 8 GPUs as the safer starting allowance. The final requirement changes with runtime, context, batch, concurrency and precision.

Which GPU Servers product should I start with for GLM-5?

Use the compatibility table to find systems that meet the recommended allowance. A system that only meets the minimum can load the model in principle but may not meet the required context or response time.

Can GLM-5 be used commercially?

The official GLM-5 checkpoint is MIT licensed. The linked official licence is authoritative; legal advice may be appropriate for the intended use and distribution route.

What a complete GLM-5 deployment needs

GPU memory is only one part of the system. Storage, data access, serving software, monitoring and administrator handover also affect a reliable deployment.

Diagram showing an approved request, a local service, an approved store and a policy-controlled data path
Diagram showing an approved request, a local service, an approved store and a policy-controlled data path
A private deployment starts with the permitted data path, access policy and logging boundary. GPU Servers technical illustration.
A private deployment starts with the permitted data path, access policy and logging boundary. GPU Servers technical illustration.
GPU server remote management dashboard with system status, access logs and sensor monitoring panels
GPU server remote management dashboard with system status, access logs and sensor monitoring panels
Supplier screenshot of the platform management interface. The final management features and access policy depend on the ordered system. OEM supplier reference image.
Supplier screenshot of the platform management interface. The final management features and access policy depend on the ordered system. OEM supplier reference image.
Diagram showing approved documents moving through a searchable index to an answer with a source citation
Diagram showing approved documents moving through a searchable index to an answer with a source citation
A retrieval workflow should connect each useful answer to approved source material and defined refusal behaviour. GPU Servers technical illustration.
A retrieval workflow should connect each useful answer to approved source material and defined refusal behaviour. GPU Servers technical illustration.
Diagram of an evidence pack containing an asset schedule, burn-in record, health readings, workload test and admin guide
Diagram of an evidence pack containing an asset schedule, burn-in record, health readings, workload test and admin guide
A complete handover includes the supplied assets, test results, operating instructions and agreed follow-up work. GPU Servers technical illustration.
A complete handover includes the supplied assets, test results, operating instructions and agreed follow-up work. GPU Servers technical illustration.