Black Forest Labs model

FLUX.2 Klein Hardware Requirements & Server Fit

Black Forest Labs documents about 13GB VRAM for FLUX.2 Klein 4B. Every GPU Servers product clears that floor; Team 32 is already suitable for one interactive image worker.

Model version
black-forest-labs/FLUX.2-klein-4B
Family and variant
FLUX.2 Klein 4B · FLUX.2-klein-4B
Source version
e7b7dc27f91d
Source updated
24 February 2026
Black Forest Labs FLUX.2 Klein 4B black-forest-labs/FLUX.2-klein-4B
Minimum GPU memory
About 13GB VRAM
Recommended hardware
24GB or more for larger images and working headroom
Licence
Apache License 2.0
Useful for
Image generation
Memory compatibility is a sizing guide. Test the exact model version and workload before choosing hardware.

Buyer verdict

Where FLUX.2 Klein 4B is a sensible fit

FLUX.2 Klein 4B is the commercially simpler FLUX.2 choice for interactive generation and editing because the 4B checkpoint is Apache-licensed, unlike the 9B non-commercial release.

These figures apply to the named model version. Quantisation, fine-tuning, context length, image resolution, batch size and serving software can materially change the hardware needed.

Hardware requirements

Minimum
About 13GB VRAM
Recommended
24GB or more for larger images and working headroom

Black Forest Labs states this requirement for the 4B checkpoint.

Resolution, reference images and batch affect peak memory.

Technical specification

FLUX.2 Klein 4B model and hardware facts

Specifications shown for source version e7b7dc27f91deacad38e78976d1f2b499d76a294, updated 24 February 2026.

Parameters
4B
Tasks
Text-to-image + image editing
Official VRAM
About 13GB
Reference output
1,024 × 1,024 in model-card example
Inference steps
4 in reference example
Licence
Apache 2.0

Product compatibility

FLUX.2 Klein 4B compatibility across all 11 GPU systems

Systems that fit without model splitting
11
Systems needing multi-GPU validation
0

A system is listed as fitting when its GPU memory meets the requirement shown above. Speed, usable context, batch size and concurrent users still need testing with the final model and software configuration.

Best larger-system route

Team 32

1 GPUs · 32GB

The smallest system in the range that reaches this model's preferred working allowance.

Explore Team 32

sme workstation

Team 32

Recommended memory route
Per GPU
32GB
Total GPU memory
32GB
GPU count
1

Also meets this model page's recommended working allowance.

Available GPU memory exceeds the model developer's published requirement. Confirm speed, context, batch size and concurrent users before purchase.

sme workstation

Company 64

Recommended memory route
Per GPU
32GB
Total GPU memory
64GB
GPU count
2

Also meets this model page's recommended working allowance.

Available GPU memory exceeds the model developer's published requirement. Confirm speed, context, batch size and concurrent users before purchase.

sme workstation

Studio 96

Recommended memory route
Per GPU
96GB
Total GPU memory
96GB
GPU count
1

Also meets this model page's recommended working allowance.

Available GPU memory exceeds the model developer's published requirement. Confirm speed, context, batch size and concurrent users before purchase.

sme workstation

Studio 192

Recommended memory route
Per GPU
96GB
Total GPU memory
192GB
GPU count
2

Also meets this model page's recommended working allowance.

Available GPU memory exceeds the model developer's published requirement. Confirm speed, context, batch size and concurrent users before purchase.

pcie rack

Value Rack 128

Recommended memory route
Per GPU
32GB
Total GPU memory
128GB
GPU count
4

Also meets this model page's recommended working allowance.

Available GPU memory exceeds the model developer's published requirement. Confirm speed, context, batch size and concurrent users before purchase.

pcie rack

Value Rack 256

Recommended memory route
Per GPU
32GB
Total GPU memory
256GB
GPU count
8

Also meets this model page's recommended working allowance.

Available GPU memory exceeds the model developer's published requirement. Confirm speed, context, batch size and concurrent users before purchase.

pcie rack

Enterprise 384

Recommended memory route
Per GPU
96GB
Total GPU memory
384GB
GPU count
4

Also meets this model page's recommended working allowance.

Available GPU memory exceeds the model developer's published requirement. Confirm speed, context, batch size and concurrent users before purchase.

pcie rack

Enterprise 768

Recommended memory route
Per GPU
96GB
Total GPU memory
768GB
GPU count
8

Also meets this model page's recommended working allowance.

Available GPU memory exceeds the model developer's published requirement. Confirm speed, context, batch size and concurrent users before purchase.

pcie rack

H200 1.1TB

Recommended memory route
Per GPU
141GB
Total GPU memory
1,128GB
GPU count
8

Also meets this model page's recommended working allowance.

Available GPU memory exceeds the model developer's published requirement. Confirm speed, context, batch size and concurrent users before purchase.

frontier partner

Frontier Native 2.3TB

Recommended memory route
Per GPU
288GB
Total GPU memory
2,304GB
GPU count
8

Also meets this model page's recommended working allowance.

Available GPU memory exceeds the model developer's published requirement. Confirm speed, context, batch size and concurrent users before purchase.

frontier partner

Frontier Rack 20TB

Recommended memory route
Per GPU
288GB
Total GPU memory
20,736GB
GPU count
72

Also meets this model page's recommended working allowance.

Available GPU memory exceeds the model developer's published requirement. Confirm speed, context, batch size and concurrent users before purchase.

Deployment reality

Strengths, limits and runtime route

What it is good at

  • Official 13GB VRAM statement.
  • Text-to-image, image editing and multi-reference input in one model.
  • Four-step distilled inference and consumer-GPU positioning.
  • Apache 2.0 commercial-use route.

Where to be cautious

  • Generated text can still be inaccurate or distorted.
  • Prompt adherence and output quality vary by subject and style.
  • The 9B sibling uses a different non-commercial licence and must not be confused with this page.
  • Safety filtering, provenance and acceptable-use controls remain deployment requirements.

Serving software

  • Diffusers
  • ComfyUI
  • FLUX.2 reference implementation

Before installation

  1. Use Diffusers or the official reference implementation and pin the checkpoint.
  2. Record resolution, steps, batch, latency and peak VRAM.
  3. Evaluate text rendering and image editing separately.
  4. Keep the 4B Apache and 9B non-commercial licences distinct in model selection.

System requirements

GPU layout
The minimum can fit on one GPU; multiple GPUs may still be useful for replicas or throughput.
Context and cache
The model context ceiling is not a guaranteed serving target. KV cache, batch size and concurrent sessions need separate capacity tests.
System RAM
Size system memory for model loading, runtime overhead, preprocessing and any CPU offload used by the final configuration.
Storage
Allow space for the pinned checkpoint, runtime images, caches, logs and at least one rollback version.
Serving software
Validate the exact checkpoint with Diffusers, ComfyUI, FLUX.2 reference implementation before acceptance.
Representative workload
Benchmark representative prompts or media at the required context, quality, latency and concurrency.

Performance: speed, usable context and concurrency depend on the selected system, software and workload. No benchmark is quoted on this page.

Commercial and legal boundary

Apache License 2.0

Commercial use: permitted

The 4B checkpoint is Apache 2.0. The 9B checkpoint is separately licensed for non-commercial use.

Always retain the applicable notices and recheck the live terms for the intended organisation, territory, use and distribution route. Obtain legal advice where required; the official licence governs use.

Read the official licence

Official sources

Technical questions

FLUX.2 Klein 4B deployment FAQ

Can Team 32 run FLUX.2 Klein 4B?

Yes. Its 32GB GPU exceeds the official approximately 13GB requirement.

Can I use FLUX.2 Klein commercially?

The 4B checkpoint is Apache 2.0. Do not assume the same for the 9B sibling, which uses a non-commercial licence.

How should image performance be tested?

Record resolution, references, steps, batch, latency, peak VRAM and an agreed quality review.

What a complete FLUX.2 Klein 4B deployment needs

GPU memory is only one part of the system. Storage, data access, serving software, monitoring and administrator handover also affect a reliable deployment.

Diagram showing an approved request, a local service, an approved store and a policy-controlled data path
Diagram showing an approved request, a local service, an approved store and a policy-controlled data path
A private deployment starts with the permitted data path, access policy and logging boundary. GPU Servers technical illustration.
A private deployment starts with the permitted data path, access policy and logging boundary. GPU Servers technical illustration.
GPU server remote management dashboard with system status, access logs and sensor monitoring panels
GPU server remote management dashboard with system status, access logs and sensor monitoring panels
Supplier screenshot of the platform management interface. The final management features and access policy depend on the ordered system. OEM supplier reference image.
Supplier screenshot of the platform management interface. The final management features and access policy depend on the ordered system. OEM supplier reference image.
Diagram showing approved documents moving through a searchable index to an answer with a source citation
Diagram showing approved documents moving through a searchable index to an answer with a source citation
A retrieval workflow should connect each useful answer to approved source material and defined refusal behaviour. GPU Servers technical illustration.
A retrieval workflow should connect each useful answer to approved source material and defined refusal behaviour. GPU Servers technical illustration.
Diagram of an evidence pack containing an asset schedule, burn-in record, health readings, workload test and admin guide
Diagram of an evidence pack containing an asset schedule, burn-in record, health readings, workload test and admin guide
A complete handover includes the supplied assets, test results, operating instructions and agreed follow-up work. GPU Servers technical illustration.
A complete handover includes the supplied assets, test results, operating instructions and agreed follow-up work. GPU Servers technical illustration.