Document and image AI

Vision and OCR models for documents and images

Compare multimodal models for document extraction, screenshots, charts, photographs and visual question answering. Image count and resolution create memory and latency costs that a text-only sizing figure cannot capture.

Selection approach

Choose for the workload, then size the complete deployment

Use OCR-specialist models for high-volume extraction and multimodal language models when interpretation and reasoning matter. Test the exact scans, layouts, handwriting and languages found in production.

01

What to compare

  • OCR fidelity versus visual reasoning depth
  • Supported image count, size and aspect ratio
  • Tables, forms, charts and handwriting requirements
  • Output schema and confidence handling

02

What changes the hardware

  • Vision encoder and image-token allocations
  • Batching across variable page sizes
  • Preprocessing CPU and storage throughput
  • Multi-page context and concurrent document queues

03

What to test before purchase

  • Field-level extraction accuracy
  • Table and reading-order preservation
  • Performance on poor scans and rotated pages
  • Latency and throughput at the intended page resolution

12 current models

Compare models by workload and GPU memory

Showing 12 models

Moonshot AI

Kimi K3

A 2.8-trillion-parameter, 104-billion-active multimodal mixture-of-experts model for long-context reasoning, coding, visual understanding and agentic work.

  • Language & reasoning
  • Coding & agents
  • Vision & OCR
Minimum GPU memory
At least 1.561TB aggregate GPU memory for the released weight files
Recommended starting system
Specialist configuration required
Licence
Kimi K3 Licence
See specifications and all 11 systems

Qwen

Qwen3.6 35B-A3B

A 36-billion-parameter multimodal mixture-of-experts model with about 3 billion active parameters, a 262,144-token default context and a strong emphasis on agentic coding.

  • Language & reasoning
  • Coding & agents
  • Vision & OCR
Minimum GPU memory
80GB-class GPU for the named BF16 weights at reduced context
Recommended starting system
Studio 96
Licence
Apache License 2.0
See specifications and all 11 systems

Mistral AI

Mistral Small 4

A 119B-total, 6B-active hybrid model that combines instruction following, reasoning, coding-agent and multimodal capabilities.

  • Language & reasoning
  • Coding & agents
  • Vision & OCR
Minimum GPU memory
256GB aggregate with supported model parallelism
Recommended starting system
Frontier Native 2.3TB
Licence
Apache License 2.0
See specifications and all 11 systems

Meta

Llama 4 Scout

Meta's 109B-total, 17B-active multimodal mixture-of-experts checkpoint with text and image input and a very long documented context.

  • Language & reasoning
  • Vision & OCR
Minimum GPU memory
At least 240GB aggregate for BF16 weights and minimal overhead
Recommended starting system
Frontier Native 2.3TB
Licence
Llama 4 Community Licence
See specifications and all 11 systems

Google

Gemma 3 27B

Google's 27B instruction-tuned multimodal model for text and image input, with a 128K context and broad language coverage.

  • Language & reasoning
  • Vision & OCR
Minimum GPU memory
64GB GPU for BF16 at a bounded context
Recommended starting system
Studio 96
Licence
Gemma Terms of Use
See specifications and all 11 systems

DeepSeek

DeepSeek-OCR 2

A 3B-class image-to-text model for document optical character recognition and layout-aware text extraction.

  • Vision & OCR
Minimum GPU memory
12GB GPU planning floor
Recommended starting system
Team 32
Licence
Apache License 2.0
See specifications and all 11 systems

Meta

Llama Guard 4

A 12B multimodal safeguard model for classifying text and image prompts and responses against Meta's hazard taxonomy.

  • Safety & moderation
  • Vision & OCR
Minimum GPU memory
28GB GPU planning floor
Recommended starting system
Team 32
Licence
Llama 4 Community Licence
See specifications and all 11 systems

Qwen

Qwen3.5 4B

A compact 4-billion-parameter multimodal model with native vision, reasoning, coding and broad multilingual support.

  • Language & reasoning
  • Coding & agents
  • Vision & OCR
Minimum GPU memory
16GB on one GPU
Recommended starting system
Team 32
Licence
Apache License 2.0
See specifications and all 11 systems

Qwen

Qwen3.5 27B

A 27-billion-parameter dense multimodal model for reasoning, coding, agents and visual understanding across 201 languages and dialects.

  • Language & reasoning
  • Coding & agents
  • Vision & OCR
Minimum GPU memory
64GB on one GPU
Recommended starting system
Studio 96
Licence
Apache License 2.0
See specifications and all 11 systems

Google DeepMind

Gemma 4 12B

Google DeepMind's 12-billion-parameter instruction-tuned Gemma 4 model, supporting text, images, video and audio input with text output.

  • Language & reasoning
  • Vision & OCR
  • Speech & audio
Minimum GPU memory
32GB on one GPU
Recommended starting system
Studio 96
Licence
Apache License 2.0
See specifications and all 11 systems

Google DeepMind

Gemma 4 31B

Google DeepMind's 30.7-billion-parameter instruction-tuned multimodal model for text and image understanding with a 256K context window.

  • Language & reasoning
  • Vision & OCR
Minimum GPU memory
64GB on one GPU
Recommended starting system
Studio 96
Licence
Apache License 2.0
See specifications and all 11 systems

StepFun

Step 3.7 Flash

A 198-billion-parameter sparse vision-language model with about 11 billion active parameters, designed for reasoning, coding, agents and visual intelligence.

  • Language & reasoning
  • Coding & agents
  • Vision & OCR
Minimum GPU memory
416GB aggregate across at least 4 GPUs
Recommended starting system
Specialist configuration required
Licence
Apache License 2.0
See specifications and all 11 systems

Model-to-hardware fit

Memory fit is the first gate, not the final recommendation

  1. 01Exact model version

    Use the precise model version, numerical format and complete software pipeline intended for production.

  2. 02Minimum memory

    Check whether it can load on one GPU or requires supported multi-GPU loading.

  3. 03Working headroom

    Allow for context, cache, batch, media encoders and concurrent users.

  4. 04Workload test

    Measure quality, latency and stability on representative work.

What a complete vision & ocr deployment needs

Model weights are only one part of the system. Data access, runtime software, storage, monitoring and administrator handover also affect a reliable deployment.

Diagram showing an approved request, a local service, an approved store and a policy-controlled data path
Diagram showing an approved request, a local service, an approved store and a policy-controlled data path
A private deployment starts with the permitted data path, access policy and logging boundary. GPU Servers technical illustration.
A private deployment starts with the permitted data path, access policy and logging boundary. GPU Servers technical illustration.
GPU server remote management dashboard with system status, access logs and sensor monitoring panels
GPU server remote management dashboard with system status, access logs and sensor monitoring panels
Supplier screenshot of the platform management interface. The final management features and access policy depend on the ordered system. OEM supplier reference image.
Supplier screenshot of the platform management interface. The final management features and access policy depend on the ordered system. OEM supplier reference image.
Diagram showing approved documents moving through a searchable index to an answer with a source citation
Diagram showing approved documents moving through a searchable index to an answer with a source citation
A retrieval workflow should connect each useful answer to approved source material and defined refusal behaviour. GPU Servers technical illustration.
A retrieval workflow should connect each useful answer to approved source material and defined refusal behaviour. GPU Servers technical illustration.
Diagram of an evidence pack containing an asset schedule, burn-in record, health readings, workload test and admin guide
Diagram of an evidence pack containing an asset schedule, burn-in record, health readings, workload test and admin guide
A complete handover includes the supplied assets, test results, operating instructions and agreed follow-up work. GPU Servers technical illustration.
A complete handover includes the supplied assets, test results, operating instructions and agreed follow-up work. GPU Servers technical illustration.