Moonshot AI model

Kimi K3 GPU Server Requirements & Compatibility

The official Kimi K3 checkpoint is a frontier-scale deployment. Its 96 safetensors files total 1.5609TB, so the current eight-H200 product does not hold the named repository checkpoint. HGX B300 and GB300 are the credible products in this range, subject to the exact runtime, topology, context and acceptance test.

Model version
moonshotai/Kimi-K3
Family and variant
Kimi K3 · Kimi-K3
Source version
9f62e4e9fffb
Source updated
27 July 2026
Moonshot AI Kimi K3 moonshotai/Kimi-K3
Minimum GPU memory
At least 1.561TB aggregate GPU memory for the released weight files
Recommended hardware
Eight-GPU HGX B300-class node or larger
Licence
Kimi K3 Licence
Useful for
Language & reasoning, Coding & agents, Vision & OCR
Memory compatibility is a sizing guide. Test the exact model version and workload before choosing hardware.

Buyer verdict

Where Kimi K3 is a sensible fit

Consider Kimi K3 when the requirement is one self-hosted model for long repository work, image-aware knowledge work or tool-using agents and the buyer accepts frontier-platform cost and operational complexity.

These figures apply to the named model version. Quantisation, fine-tuning, context length, image resolution, batch size and serving software can materially change the hardware needed.

Hardware requirements

Minimum
At least 1.561TB aggregate GPU memory for the released weight files
Recommended
Eight-GPU HGX B300-class node or larger

This is a file-size floor only. Runtime overhead, KV cache and production headroom increase the requirement.

Runtime documentation includes B300 deployment recipes. GPU Servers has not measured this checkpoint.

Technical specification

Kimi K3 model and hardware facts

Specifications shown for source version 9f62e4e9fffbd0a83ddd60e1c209d828994b3569, updated 27 July 2026.

Architecture
Sparse mixture of experts
2.8T total, 104B active
Context
1,048,576 tokens
Model ceiling, not a server guarantee
Modalities
Text + image → text
Native format
MXFP4 weights / MXFP8 activations
Repository weights
1,560.94GB
96 safetensors files checked 30 July 2026
Serving
vLLM, SGLang, TokenSpeed

Product compatibility

Kimi K3 compatibility across all 11 GPU systems

Systems that fit without model splitting
0
Systems needing multi-GPU validation
2

A system is listed as fitting when its GPU memory meets the requirement shown above. Speed, usable context, batch size and concurrent users still need testing with the final model and software configuration.

First multi-GPU opportunity

Frontier Native 2.3TB

8 GPUs · 2,304GB

The first system with enough installed GPU memory when supported software divides the model across its GPUs. We confirm software support, GPU layout, context and performance during sizing.

Explore Frontier Native 2.3TB

sme workstation

Team 32

Larger-memory system recommended
Per GPU
32GB
Total GPU memory
32GB
GPU count
1

This system provides 32GB per GPU and 32GB in total. The model requires 1561GB aggregate across at least 8 GPUs.

For this model, explore Frontier Native 2.3TB .

sme workstation

Company 64

Larger-memory system recommended
Per GPU
32GB
Total GPU memory
64GB
GPU count
2

This system provides 32GB per GPU and 64GB in total. The model requires 1561GB aggregate across at least 8 GPUs.

For this model, explore Frontier Native 2.3TB .

sme workstation

Studio 96

Larger-memory system recommended
Per GPU
96GB
Total GPU memory
96GB
GPU count
1

This system provides 96GB per GPU and 96GB in total. The model requires 1561GB aggregate across at least 8 GPUs.

For this model, explore Frontier Native 2.3TB .

sme workstation

Studio 192

Larger-memory system recommended
Per GPU
96GB
Total GPU memory
192GB
GPU count
2

This system provides 96GB per GPU and 192GB in total. The model requires 1561GB aggregate across at least 8 GPUs.

For this model, explore Frontier Native 2.3TB .

pcie rack

Value Rack 128

Larger-memory system recommended
Per GPU
32GB
Total GPU memory
128GB
GPU count
4

This system provides 32GB per GPU and 128GB in total. The model requires 1561GB aggregate across at least 8 GPUs.

For this model, explore Frontier Native 2.3TB .

pcie rack

Value Rack 256

Larger-memory system recommended
Per GPU
32GB
Total GPU memory
256GB
GPU count
8

This system provides 32GB per GPU and 256GB in total. The model requires 1561GB aggregate across at least 8 GPUs.

For this model, explore Frontier Native 2.3TB .

pcie rack

Enterprise 384

Larger-memory system recommended
Per GPU
96GB
Total GPU memory
384GB
GPU count
4

This system provides 96GB per GPU and 384GB in total. The model requires 1561GB aggregate across at least 8 GPUs.

For this model, explore Frontier Native 2.3TB .

pcie rack

Enterprise 768

Larger-memory system recommended
Per GPU
96GB
Total GPU memory
768GB
GPU count
8

This system provides 96GB per GPU and 768GB in total. The model requires 1561GB aggregate across at least 8 GPUs.

For this model, explore Frontier Native 2.3TB .

pcie rack

H200 1.1TB

Larger-memory system recommended
Per GPU
141GB
Total GPU memory
1,128GB
GPU count
8

This system provides 141GB per GPU and 1128GB in total. The model requires 1561GB aggregate across at least 8 GPUs.

For this model, explore Frontier Native 2.3TB .

frontier partner

Frontier Native 2.3TB

Multi-GPU route available
Per GPU
288GB
Total GPU memory
2,304GB
GPU count
8

Total installed memory is sufficient, but the model must be divided across GPUs. Confirm that the serving software supports this GPU layout and test the required context and speed.

frontier partner

Frontier Rack 20TB

Multi-GPU route available
Per GPU
288GB
Total GPU memory
20,736GB
GPU count
72

Total installed memory is sufficient, but the model must be divided across GPUs. Confirm that the serving software supports this GPU layout and test the required context and speed.

Deployment reality

Strengths, limits and runtime route

What it is good at

  • Native text and image input with a documented 1,048,576-token context limit.
  • Kimi Delta Attention, gated MLA and a sparse 896-expert architecture with 16 selected experts per token.
  • Official deployment routes for vLLM, SGLang and TokenSpeed.
  • Quantisation-aware MXFP4 weights with MXFP8 activations rather than an unofficial post-release conversion.

Where to be cautious

  • The repository weight total alone does not prove loaded memory, usable one-million-token context, latency or concurrency.
  • Eight H200 GPUs provide about 1.128TB aggregate HBM, below the current official weight-file total.
  • The Kimi K3 licence is not Apache or MIT. Commercial and redistribution terms must be reviewed for the intended deployment.
  • The exact released checkpoint requires recent runtime support and a multi-GPU topology validated against Moonshot's serving format.

Serving software

  • vLLM
  • SGLang
  • TokenSpeed
  • Transformers

Before installation

  1. Pin commit 9f62e4e9fffbd0a83ddd60e1c209d828994b3569 before staging or acceptance.
  2. Record the vLLM or SGLang version, container digest, tensor and pipeline parallel settings and the actual context target.
  3. Treat the SGLang B300 and multi-node H200 recipes as vendor/runtime documentation, not a benchmark on a GPU Servers product.
  4. Preserve reasoning_content and tool calls between turns as required by Moonshot's multi-turn usage guidance.

System requirements

GPU layout
Plan for at least 8 GPUs and validate model parallelism on the final topology.
Context and cache
The model context ceiling is not a guaranteed serving target. KV cache, batch size and concurrent sessions need separate capacity tests.
System RAM
Size system memory for model loading, runtime overhead, preprocessing and any CPU offload used by the final configuration.
Storage
Allow space for the pinned checkpoint, runtime images, caches, logs and at least one rollback version.
Serving software
Validate the exact checkpoint with vLLM, SGLang, TokenSpeed, Transformers before acceptance.
Representative workload
Benchmark representative prompts or media at the required context, quality, latency and concurrency.

Performance: speed, usable context and concurrency depend on the selected system, software and workload. No benchmark is quoted on this page.

Commercial and legal boundary

Kimi K3 Licence

Commercial use: conditional

Use is governed by Moonshot's model-specific licence. Review the current text and intended distribution model before commercial deployment.

Always retain the applicable notices and recheck the live terms for the intended organisation, territory, use and distribution route. Obtain legal advice where required; the official licence governs use.

Read the official licence

Official sources

Technical questions

Kimi K3 deployment FAQ

Can an eight-GPU H200 server run the official Kimi K3 checkpoint?

Not as a demonstrated native-checkpoint fit. Eight H200 GPUs provide less aggregate memory than the current official weight-file total. A different conversion or multi-node recipe needs its own checkpoint, quality and runtime evidence.

Is HGX B300 enough for Kimi K3?

Its 2.304TB aggregate HBM clears the released weight-file total and runtime documentation includes a B300 recipe. This supports memory feasibility, but does not establish throughput or full-context performance.

Does Kimi K3 support images and video?

Moonshot documents text and image as the model's formal modalities and describes video understanding in its release material. The exact video preprocessing, token budget and runtime behaviour still need a workload test.

What a complete Kimi K3 deployment needs

GPU memory is only one part of the system. Storage, data access, serving software, monitoring and administrator handover also affect a reliable deployment.

Diagram showing an approved request, a local service, an approved store and a policy-controlled data path
Diagram showing an approved request, a local service, an approved store and a policy-controlled data path
A private deployment starts with the permitted data path, access policy and logging boundary. GPU Servers technical illustration.
A private deployment starts with the permitted data path, access policy and logging boundary. GPU Servers technical illustration.
GPU server remote management dashboard with system status, access logs and sensor monitoring panels
GPU server remote management dashboard with system status, access logs and sensor monitoring panels
Supplier screenshot of the platform management interface. The final management features and access policy depend on the ordered system. OEM supplier reference image.
Supplier screenshot of the platform management interface. The final management features and access policy depend on the ordered system. OEM supplier reference image.
Diagram showing approved documents moving through a searchable index to an answer with a source citation
Diagram showing approved documents moving through a searchable index to an answer with a source citation
A retrieval workflow should connect each useful answer to approved source material and defined refusal behaviour. GPU Servers technical illustration.
A retrieval workflow should connect each useful answer to approved source material and defined refusal behaviour. GPU Servers technical illustration.
Diagram of an evidence pack containing an asset schedule, burn-in record, health readings, workload test and admin guide
Diagram of an evidence pack containing an asset schedule, burn-in record, health readings, workload test and admin guide
A complete handover includes the supplied assets, test results, operating instructions and agreed follow-up work. GPU Servers technical illustration.
A complete handover includes the supplied assets, test results, operating instructions and agreed follow-up work. GPU Servers technical illustration.