DeepSeek model

DeepSeek-OCR 2 Hardware Requirements & Compatibility

The BF16 weight file is about 6.8GB. Every current GPU Servers product clears the model-weight floor; useful product choice is driven by page throughput, resolution, queue size, preprocessing and whether OCR shares a GPU with another service.

Model version
deepseek-ai/DeepSeek-OCR-2
Family and variant
DeepSeek-OCR 2 · DeepSeek-OCR-2
Source version
aaa02f381194
Source updated
3 February 2026
DeepSeek DeepSeek-OCR 2 deepseek-ai/DeepSeek-OCR-2
Minimum GPU memory
12GB GPU planning floor
Recommended hardware
24GB or more for practical document batches
Licence
Apache License 2.0
Useful for
Vision & OCR
Memory compatibility is a sizing guide. Test the exact model version and workload before choosing hardware.

Buyer verdict

Where DeepSeek-OCR 2 is a sensible fit

DeepSeek-OCR 2 is relevant for scanned documents, page images, structured extraction and OCR pipelines that need a dedicated local vision model under Apache 2.0.

These figures apply to the named model version. Quantisation, fine-tuning, context length, image resolution, batch size and serving software can materially change the hardware needed.

Hardware requirements

Minimum
12GB GPU planning floor
Recommended
24GB or more for practical document batches

The official weights are 6.8GB; image tensors, runtime and batch need additional memory.

Throughput still depends on resolution, preprocessing, batch and storage.

Technical specification

DeepSeek-OCR 2 model and hardware facts

Specifications shown for source version aaa02f3811945a91062062994c5c4a3f4c0af2b0, updated 3 February 2026.

Parameters
About 3.4B
Modality
Image + prompt → text
Repository weights
6.78GB BF16
Primary workload
Document OCR
Library
Transformers with model-specific code
Licence
Apache 2.0

Product compatibility

DeepSeek-OCR 2 compatibility across all 11 GPU systems

Systems that fit without model splitting
11
Systems needing multi-GPU validation
0

A system is listed as fitting when its GPU memory meets the requirement shown above. Speed, usable context, batch size and concurrent users still need testing with the final model and software configuration.

Best larger-system route

Team 32

1 GPUs · 32GB

The smallest system in the range that reaches this model's preferred working allowance.

Explore Team 32

sme workstation

Team 32

Recommended memory route
Per GPU
32GB
Total GPU memory
32GB
GPU count
1

Also meets this model page's recommended working allowance.

Available GPU memory exceeds the calculated or estimated requirement. Confirm the final precision, software, workload and performance before purchase.

sme workstation

Company 64

Recommended memory route
Per GPU
32GB
Total GPU memory
64GB
GPU count
2

Also meets this model page's recommended working allowance.

Available GPU memory exceeds the calculated or estimated requirement. Confirm the final precision, software, workload and performance before purchase.

sme workstation

Studio 96

Recommended memory route
Per GPU
96GB
Total GPU memory
96GB
GPU count
1

Also meets this model page's recommended working allowance.

Available GPU memory exceeds the calculated or estimated requirement. Confirm the final precision, software, workload and performance before purchase.

sme workstation

Studio 192

Recommended memory route
Per GPU
96GB
Total GPU memory
192GB
GPU count
2

Also meets this model page's recommended working allowance.

Available GPU memory exceeds the calculated or estimated requirement. Confirm the final precision, software, workload and performance before purchase.

pcie rack

Value Rack 128

Recommended memory route
Per GPU
32GB
Total GPU memory
128GB
GPU count
4

Also meets this model page's recommended working allowance.

Available GPU memory exceeds the calculated or estimated requirement. Confirm the final precision, software, workload and performance before purchase.

pcie rack

Value Rack 256

Recommended memory route
Per GPU
32GB
Total GPU memory
256GB
GPU count
8

Also meets this model page's recommended working allowance.

Available GPU memory exceeds the calculated or estimated requirement. Confirm the final precision, software, workload and performance before purchase.

pcie rack

Enterprise 384

Recommended memory route
Per GPU
96GB
Total GPU memory
384GB
GPU count
4

Also meets this model page's recommended working allowance.

Available GPU memory exceeds the calculated or estimated requirement. Confirm the final precision, software, workload and performance before purchase.

pcie rack

Enterprise 768

Recommended memory route
Per GPU
96GB
Total GPU memory
768GB
GPU count
8

Also meets this model page's recommended working allowance.

Available GPU memory exceeds the calculated or estimated requirement. Confirm the final precision, software, workload and performance before purchase.

pcie rack

H200 1.1TB

Recommended memory route
Per GPU
141GB
Total GPU memory
1,128GB
GPU count
8

Also meets this model page's recommended working allowance.

Available GPU memory exceeds the calculated or estimated requirement. Confirm the final precision, software, workload and performance before purchase.

frontier partner

Frontier Native 2.3TB

Recommended memory route
Per GPU
288GB
Total GPU memory
2,304GB
GPU count
8

Also meets this model page's recommended working allowance.

Available GPU memory exceeds the calculated or estimated requirement. Confirm the final precision, software, workload and performance before purchase.

frontier partner

Frontier Rack 20TB

Recommended memory route
Per GPU
288GB
Total GPU memory
20,736GB
GPU count
72

Also meets this model page's recommended working allowance.

Available GPU memory exceeds the calculated or estimated requirement. Confirm the final precision, software, workload and performance before purchase.

Deployment reality

Strengths, limits and runtime route

What it is good at

  • Compact model relative to general multimodal LLMs.
  • Apache 2.0 licence and public official checkpoint.
  • Suitable as one stage in a local document-ingestion pipeline.
  • Leaves room on a 32GB GPU for preprocessing and bounded parallel work.

Where to be cautious

  • OCR accuracy varies with scan quality, language, handwriting, tables and page layout.
  • A model loading successfully does not establish pages per minute or extraction quality.
  • The official code uses custom model support and version requirements that need pinning.
  • Document retention, access and human review remain part of the deployment.

Serving software

  • Transformers
  • DeepSeek reference code

Before installation

  1. Benchmark with a labelled sample of the buyer's invoices, contracts, forms or technical documents.
  2. Record input resolution, page count, batch, latency and peak VRAM.
  3. Keep OCR output and downstream extraction evaluation separate.
  4. A 32GB workstation is often enough for a pilot; rack products matter when several workers or large ingestion queues are required.

System requirements

GPU layout
The minimum can fit on one GPU; multiple GPUs may still be useful for replicas or throughput.
Context and cache
The model context ceiling is not a guaranteed serving target. KV cache, batch size and concurrent sessions need separate capacity tests.
System RAM
Size system memory for model loading, runtime overhead, preprocessing and any CPU offload used by the final configuration.
Storage
Allow space for the pinned checkpoint, runtime images, caches, logs and at least one rollback version.
Serving software
Validate the exact checkpoint with Transformers, DeepSeek reference code before acceptance.
Representative workload
Benchmark representative prompts or media at the required context, quality, latency and concurrency.

Performance: speed, usable context and concurrency depend on the selected system, software and workload. No benchmark is quoted on this page.

Commercial and legal boundary

Apache License 2.0

Commercial use: permitted

Apache 2.0 permits commercial use. Document rights and personal-data handling remain separate.

Always retain the applicable notices and recheck the live terms for the intended organisation, territory, use and distribution route. Obtain legal advice where required; the official licence governs use.

Read the official licence

Official sources

Technical questions

DeepSeek-OCR 2 deployment FAQ

Does every GPU Servers product fit DeepSeek-OCR 2?

All current products exceed the planning memory floor. The right size depends on page resolution, throughput, queueing and any language model used after OCR.

Is OCR accuracy guaranteed?

No. Accuracy must be measured against representative pages and an agreed field or character-level acceptance method.

Do I need a rack server for OCR?

Not for a small pilot. Rack systems become relevant for multiple workers, sustained ingestion or a combined OCR, embedding and language-model pipeline.

What a complete DeepSeek-OCR 2 deployment needs

GPU memory is only one part of the system. Storage, data access, serving software, monitoring and administrator handover also affect a reliable deployment.

Diagram showing an approved request, a local service, an approved store and a policy-controlled data path
Diagram showing an approved request, a local service, an approved store and a policy-controlled data path
A private deployment starts with the permitted data path, access policy and logging boundary. GPU Servers technical illustration.
A private deployment starts with the permitted data path, access policy and logging boundary. GPU Servers technical illustration.
GPU server remote management dashboard with system status, access logs and sensor monitoring panels
GPU server remote management dashboard with system status, access logs and sensor monitoring panels
Supplier screenshot of the platform management interface. The final management features and access policy depend on the ordered system. OEM supplier reference image.
Supplier screenshot of the platform management interface. The final management features and access policy depend on the ordered system. OEM supplier reference image.
Diagram showing approved documents moving through a searchable index to an answer with a source citation
Diagram showing approved documents moving through a searchable index to an answer with a source citation
A retrieval workflow should connect each useful answer to approved source material and defined refusal behaviour. GPU Servers technical illustration.
A retrieval workflow should connect each useful answer to approved source material and defined refusal behaviour. GPU Servers technical illustration.
Diagram of an evidence pack containing an asset schedule, burn-in record, health readings, workload test and admin guide
Diagram of an evidence pack containing an asset schedule, burn-in record, health readings, workload test and admin guide
A complete handover includes the supplied assets, test results, operating instructions and agreed follow-up work. GPU Servers technical illustration.
A complete handover includes the supplied assets, test results, operating instructions and agreed follow-up work. GPU Servers technical illustration.