Independent deployment field guide

Choose the Best Open-Source AI Model for Business Workloads

Compare current models for language, coding, vision, OCR, embeddings, speech, image, video, 3D and safety. See the minimum GPU memory, licence and compatibility across all eleven GPU Servers products.

Start with the decision

The best model is the smallest one that passes your workload.

A useful choice depends on task quality, licence, runtime maturity, context, concurrency and deployment cost. Parameter count is not a substitute for testing the work your organisation actually performs.

A best open-source AI model for business shortlist may include both open-source and source-available open-weight releases because buyers encounter both. Each page states the applicable licence rather than treating “open” as one legal category.

Browse by task

Find models for the work you need to run

One model can appear in more than one category when it handles several types of work.

33 current models

Compare models by workload and GPU memory

Showing 33 models

Moonshot AI

Kimi K3

A 2.8-trillion-parameter, 104-billion-active multimodal mixture-of-experts model for long-context reasoning, coding, visual understanding and agentic work.

  • Language & reasoning
  • Coding & agents
  • Vision & OCR
Minimum GPU memory
At least 1.561TB aggregate GPU memory for the released weight files
Recommended starting system
Specialist configuration required
Licence
Kimi K3 Licence
See specifications and all 11 systems

Qwen

Qwen3.6 35B-A3B

A 36-billion-parameter multimodal mixture-of-experts model with about 3 billion active parameters, a 262,144-token default context and a strong emphasis on agentic coding.

  • Language & reasoning
  • Coding & agents
  • Vision & OCR
Minimum GPU memory
80GB-class GPU for the named BF16 weights at reduced context
Recommended starting system
Studio 96
Licence
Apache License 2.0
See specifications and all 11 systems

DeepSeek

DeepSeek V4 Flash

A 284B-total, 13B-active mixture-of-experts language model with one-million-token context and a mixed FP4/FP8 release format.

  • Language & reasoning
  • Coding & agents
Minimum GPU memory
At least 176GB aggregate GPU memory with a supported sharding route
Recommended starting system
Frontier Native 2.3TB
Licence
MIT License
See specifications and all 11 systems

Mistral AI

Mistral Small 4

A 119B-total, 6B-active hybrid model that combines instruction following, reasoning, coding-agent and multimodal capabilities.

  • Language & reasoning
  • Coding & agents
  • Vision & OCR
Minimum GPU memory
256GB aggregate with supported model parallelism
Recommended starting system
Frontier Native 2.3TB
Licence
Apache License 2.0
See specifications and all 11 systems

Meta

Llama 4 Scout

Meta's 109B-total, 17B-active multimodal mixture-of-experts checkpoint with text and image input and a very long documented context.

  • Language & reasoning
  • Vision & OCR
Minimum GPU memory
At least 240GB aggregate for BF16 weights and minimal overhead
Recommended starting system
Frontier Native 2.3TB
Licence
Llama 4 Community Licence
See specifications and all 11 systems

OpenAI

GPT-OSS 120B

OpenAI's 117B-total, 5.1B-active open-weight reasoning and agentic model, released with native MXFP4 MoE weights.

  • Language & reasoning
  • Coding & agents
Minimum GPU memory
One 80GB GPU
Recommended starting system
Studio 96
Licence
Apache License 2.0
See specifications and all 11 systems

Google

Gemma 3 27B

Google's 27B instruction-tuned multimodal model for text and image input, with a 128K context and broad language coverage.

  • Language & reasoning
  • Vision & OCR
Minimum GPU memory
64GB GPU for BF16 at a bounded context
Recommended starting system
Studio 96
Licence
Gemma Terms of Use
See specifications and all 11 systems

DeepSeek

DeepSeek-OCR 2

A 3B-class image-to-text model for document optical character recognition and layout-aware text extraction.

  • Vision & OCR
Minimum GPU memory
12GB GPU planning floor
Recommended starting system
Team 32
Licence
Apache License 2.0
See specifications and all 11 systems

Qwen

Qwen3-Embedding 8B

The largest Qwen3 embedding checkpoint, designed for multilingual dense retrieval, classification, clustering and text matching.

  • Embeddings & RAG
Minimum GPU memory
20GB GPU planning floor
Recommended starting system
Team 32
Licence
Apache License 2.0
See specifications and all 11 systems

BAAI

BGE-M3

A multilingual embedding model that supports dense, sparse and multi-vector retrieval with 1,024 dimensions and inputs up to 8,192 tokens.

  • Embeddings & RAG
Minimum GPU memory
4GB GPU planning floor, or CPU for light use
Recommended starting system
Team 32
Licence
MIT License
See specifications and all 11 systems

Qwen

Qwen3-ASR 1.7B

A compact automatic speech recognition checkpoint released in July 2026 for multilingual transcription and audio understanding.

  • Speech & audio
Minimum GPU memory
8GB GPU planning floor
Recommended starting system
Team 32
Licence
Apache License 2.0
See specifications and all 11 systems

OpenAI

Whisper Large V3

OpenAI's 1.55B-parameter multilingual speech-recognition and translation checkpoint, widely supported across transcription runtimes.

  • Speech & audio
Minimum GPU memory
6GB GPU planning floor for an optimised inference runtime
Recommended starting system
Team 32
Licence
Apache License 2.0
See specifications and all 11 systems

Black Forest Labs

FLUX.2 Klein 4B

A compact rectified-flow model for text-to-image, image editing and multi-reference work, released under Apache 2.0.

  • Image generation
Minimum GPU memory
About 13GB VRAM
Recommended starting system
Team 32
Licence
Apache License 2.0
See specifications and all 11 systems

Qwen

Qwen-Image 2512

The December 2025 Qwen-Image update for text-to-image generation, with improved realism, natural detail and text rendering.

  • Image generation
Minimum GPU memory
64GB GPU as a tight full-weight floor
Recommended starting system
Studio 96
Licence
Apache License 2.0
See specifications and all 11 systems

Wan AI

Wan2.2 TI2V 5B

A 5B text-and-image-to-video model that supports 720p generation and an official single-GPU offload route.

  • Video generation
Minimum GPU memory
24GB VRAM with documented CPU offload settings
Recommended starting system
Studio 96
Licence
Apache License 2.0
See specifications and all 11 systems

Tencent

HunyuanVideo 1.5

An 8.3B text-to-video and image-to-video model with 480p and 720p checkpoints, optional super-resolution and distilled workflows.

  • Video generation
Minimum GPU memory
14GB VRAM with model offloading
Recommended starting system
Team 32
Licence
Tencent Hunyuan Community Licence
See specifications and all 11 systems

Tencent

Hunyuan3D 2.1

Tencent's image-to-3D generation pipeline for shape and texture creation, with open model weights and a model-specific community licence.

  • 3D generation
Minimum GPU memory
24GB single-GPU planning floor
Recommended starting system
Studio 96
Licence
Tencent Hunyuan Community Licence
See specifications and all 11 systems

Meta

Llama Guard 4

A 12B multimodal safeguard model for classifying text and image prompts and responses against Meta's hazard taxonomy.

  • Safety & moderation
  • Vision & OCR
Minimum GPU memory
28GB GPU planning floor
Recommended starting system
Team 32
Licence
Llama 4 Community Licence
See specifications and all 11 systems

Stability AI

Stable Diffusion 3.5 Large

An 8B-class text-to-image model in the Stable Diffusion 3.5 family, with a large ecosystem of Diffusers and ComfyUI workflows.

  • Image generation
Minimum GPU memory
24GB planning floor with an optimised or offload workflow
Recommended starting system
Team 32
Licence
Stability AI Community Licence
See specifications and all 11 systems

OpenAI

GPT-OSS 20B

A 21-billion-parameter sparse reasoning model with 3.6 billion active parameters, native tool use and an MXFP4 release intended for local or specialised work.

  • Language & reasoning
  • Coding & agents
Minimum GPU memory
16GB on one GPU
Recommended starting system
Team 32
Licence
Apache License 2.0
See specifications and all 11 systems

Qwen

Qwen3-Coder-Next

An 80-billion-total, 3-billion-active sparse model designed for coding agents, local development and long repository context.

  • Coding & agents
  • Language & reasoning
Minimum GPU memory
160GB aggregate across at least 2 GPUs
Recommended starting system
Frontier Native 2.3TB
Licence
Apache License 2.0
See specifications and all 11 systems

Qwen

Qwen3.5 4B

A compact 4-billion-parameter multimodal model with native vision, reasoning, coding and broad multilingual support.

  • Language & reasoning
  • Coding & agents
  • Vision & OCR
Minimum GPU memory
16GB on one GPU
Recommended starting system
Team 32
Licence
Apache License 2.0
See specifications and all 11 systems

Qwen

Qwen3.5 27B

A 27-billion-parameter dense multimodal model for reasoning, coding, agents and visual understanding across 201 languages and dialects.

  • Language & reasoning
  • Coding & agents
  • Vision & OCR
Minimum GPU memory
64GB on one GPU
Recommended starting system
Studio 96
Licence
Apache License 2.0
See specifications and all 11 systems

Google DeepMind

Gemma 4 12B

Google DeepMind's 12-billion-parameter instruction-tuned Gemma 4 model, supporting text, images, video and audio input with text output.

  • Language & reasoning
  • Vision & OCR
  • Speech & audio
Minimum GPU memory
32GB on one GPU
Recommended starting system
Studio 96
Licence
Apache License 2.0
See specifications and all 11 systems

Google DeepMind

Gemma 4 31B

Google DeepMind's 30.7-billion-parameter instruction-tuned multimodal model for text and image understanding with a 256K context window.

  • Language & reasoning
  • Vision & OCR
Minimum GPU memory
64GB on one GPU
Recommended starting system
Studio 96
Licence
Apache License 2.0
See specifications and all 11 systems

StepFun

Step 3.7 Flash

A 198-billion-parameter sparse vision-language model with about 11 billion active parameters, designed for reasoning, coding, agents and visual intelligence.

  • Language & reasoning
  • Coding & agents
  • Vision & OCR
Minimum GPU memory
416GB aggregate across at least 4 GPUs
Recommended starting system
Specialist configuration required
Licence
Apache License 2.0
See specifications and all 11 systems

Z.ai

GLM-5

A 744-billion-total, 40-billion-active sparse model aimed at complex systems engineering and long-horizon agentic tasks.

  • Language & reasoning
  • Coding & agents
Minimum GPU memory
1536GB aggregate across at least 8 GPUs
Recommended starting system
Specialist configuration required
Licence
MIT License
See specifications and all 11 systems

MiniMax

MiniMax M2.5

A sparse agentic model trained for coding, tool use, search and office work across more than ten programming languages.

  • Language & reasoning
  • Coding & agents
Minimum GPU memory
256GB aggregate across at least 2 GPUs
Recommended starting system
Frontier Native 2.3TB
Licence
Modified MIT License
See specifications and all 11 systems

Qwen

Qwen3-TTS 1.7B

A multilingual text-to-speech model with streaming and non-streaming generation, instruction-controlled delivery and nine supplied voice timbres.

  • Speech & audio
Minimum GPU memory
8GB on one GPU
Recommended starting system
Team 32
Licence
Apache License 2.0
See specifications and all 11 systems

Qwen

Qwen3-Reranker 8B

An 8-billion-parameter multilingual reranker for scoring retrieved passages across more than 100 natural and programming languages.

  • Embeddings & RAG
Minimum GPU memory
24GB on one GPU
Recommended starting system
Team 32
Licence
Apache License 2.0
See specifications and all 11 systems

Black Forest Labs

FLUX.2 Klein 9B

A 9-billion-parameter rectified-flow image model for text-to-image generation and multi-reference image editing.

  • Image generation
Minimum GPU memory
64GB on one GPU
Recommended starting system
Studio 96
Licence
FLUX Non-Commercial License
See specifications and all 11 systems

StepFun

Step-Audio 2 Mini

An end-to-end audio-language model for spoken conversation, audio understanding, paralinguistic cues and audio tool use.

  • Speech & audio
  • Language & reasoning
Minimum GPU memory
24GB on one GPU
Recommended starting system
Team 32
Licence
Apache License 2.0
See specifications and all 11 systems

Understanding compatibility

How each model is matched to a server

GPU memory provides the first useful sizing check. The final result also depends on model precision, serving software, context, batch, concurrent users and the way multiple GPUs are connected.

Recommended memory route
The server reaches the model's preferred working allowance. Speed and usable context are confirmed during sizing.
Minimum memory route
The server clears the lower memory requirement, with a larger system advised when context, batch or concurrency matter.
Multi-GPU opportunity
Total installed memory is sufficient when supported software can divide the model across the system's GPUs.
Larger-memory system recommended
The model page identifies the smallest larger-system route that best meets its GPU-memory needs.

Model choice connects software evidence to physical infrastructure

Memory, runtime, data boundaries and service access affect whether a model can be deployed reliably. Test the exact model version with your real workload before choosing hardware.

Diagram showing an approved request, a local service, an approved store and a policy-controlled data path
Diagram showing an approved request, a local service, an approved store and a policy-controlled data path
A private deployment starts with the permitted data path, access policy and logging boundary. GPU Servers technical illustration.
A private deployment starts with the permitted data path, access policy and logging boundary. GPU Servers technical illustration.
GPU server remote management dashboard with system status, access logs and sensor monitoring panels
GPU server remote management dashboard with system status, access logs and sensor monitoring panels
Supplier screenshot of the platform management interface. The final management features and access policy depend on the ordered system. OEM supplier reference image.
Supplier screenshot of the platform management interface. The final management features and access policy depend on the ordered system. OEM supplier reference image.
Diagram showing approved documents moving through a searchable index to an answer with a source citation
Diagram showing approved documents moving through a searchable index to an answer with a source citation
A retrieval workflow should connect each useful answer to approved source material and defined refusal behaviour. GPU Servers technical illustration.
A retrieval workflow should connect each useful answer to approved source material and defined refusal behaviour. GPU Servers technical illustration.
Diagram of an evidence pack containing an asset schedule, burn-in record, health readings, workload test and admin guide
Diagram of an evidence pack containing an asset schedule, burn-in record, health readings, workload test and admin guide
A complete handover includes the supplied assets, test results, operating instructions and agreed follow-up work. GPU Servers technical illustration.
A complete handover includes the supplied assets, test results, operating instructions and agreed follow-up work. GPU Servers technical illustration.

Need another model?

Ask us to check a model that is not listed

Send us the Hugging Face, GitHub or official model link. We will compare the exact version, licence and memory requirement with the complete server range.

A similar name is not enough to establish compatibility. Quantisation, context length, image resolution and runtime can materially change the hardware requirement.

Before choosing

  • Check that the licence permits your intended use
  • Size for context, cache, batch and concurrent users
  • Use a runtime that supports the exact model version
  • Test representative work before accepting the hardware
Ask us to check a model