Private language models

Language and reasoning models for private AI

Compare self-hosted language models for assistants, analysis, long documents, structured output and tool use. Hardware fit starts with the exact checkpoint and precision, then expands for context, cache and concurrent users.

Selection approach

Choose for the workload, then size the complete deployment

Start with the smallest model that passes your own accuracy and reasoning tests. Larger total parameter counts can improve capability, but sparse activation, runtime maturity and memory format materially change cost and speed.

01

What to compare

  • Quality on your real prompts and document formats
  • Native context and the context actually required
  • Tool calling, structured output and multilingual needs
  • Licence terms for commercial use and redistribution

02

What changes the hardware

  • Weight memory at the chosen precision
  • KV-cache growth with context and concurrency
  • Single-GPU simplicity versus multi-GPU model parallelism
  • CPU RAM and storage for staging, offload and rollback

03

What to test before purchase

  • Task success against a held-out business dataset
  • Time to first token and tokens per second
  • Maximum useful context before quality or speed degrades
  • Concurrent-session stability and recovery behaviour

17 current models

Compare models by workload and GPU memory

Showing 17 models

Moonshot AI

Kimi K3

A 2.8-trillion-parameter, 104-billion-active multimodal mixture-of-experts model for long-context reasoning, coding, visual understanding and agentic work.

  • Language & reasoning
  • Coding & agents
  • Vision & OCR
Minimum GPU memory
At least 1.561TB aggregate GPU memory for the released weight files
Recommended starting system
Specialist configuration required
Licence
Kimi K3 Licence
See specifications and all 11 systems

Qwen

Qwen3.6 35B-A3B

A 36-billion-parameter multimodal mixture-of-experts model with about 3 billion active parameters, a 262,144-token default context and a strong emphasis on agentic coding.

  • Language & reasoning
  • Coding & agents
  • Vision & OCR
Minimum GPU memory
80GB-class GPU for the named BF16 weights at reduced context
Recommended starting system
Studio 96
Licence
Apache License 2.0
See specifications and all 11 systems

DeepSeek

DeepSeek V4 Flash

A 284B-total, 13B-active mixture-of-experts language model with one-million-token context and a mixed FP4/FP8 release format.

  • Language & reasoning
  • Coding & agents
Minimum GPU memory
At least 176GB aggregate GPU memory with a supported sharding route
Recommended starting system
Frontier Native 2.3TB
Licence
MIT License
See specifications and all 11 systems

Mistral AI

Mistral Small 4

A 119B-total, 6B-active hybrid model that combines instruction following, reasoning, coding-agent and multimodal capabilities.

  • Language & reasoning
  • Coding & agents
  • Vision & OCR
Minimum GPU memory
256GB aggregate with supported model parallelism
Recommended starting system
Frontier Native 2.3TB
Licence
Apache License 2.0
See specifications and all 11 systems

Meta

Llama 4 Scout

Meta's 109B-total, 17B-active multimodal mixture-of-experts checkpoint with text and image input and a very long documented context.

  • Language & reasoning
  • Vision & OCR
Minimum GPU memory
At least 240GB aggregate for BF16 weights and minimal overhead
Recommended starting system
Frontier Native 2.3TB
Licence
Llama 4 Community Licence
See specifications and all 11 systems

OpenAI

GPT-OSS 120B

OpenAI's 117B-total, 5.1B-active open-weight reasoning and agentic model, released with native MXFP4 MoE weights.

  • Language & reasoning
  • Coding & agents
Minimum GPU memory
One 80GB GPU
Recommended starting system
Studio 96
Licence
Apache License 2.0
See specifications and all 11 systems

Google

Gemma 3 27B

Google's 27B instruction-tuned multimodal model for text and image input, with a 128K context and broad language coverage.

  • Language & reasoning
  • Vision & OCR
Minimum GPU memory
64GB GPU for BF16 at a bounded context
Recommended starting system
Studio 96
Licence
Gemma Terms of Use
See specifications and all 11 systems

OpenAI

GPT-OSS 20B

A 21-billion-parameter sparse reasoning model with 3.6 billion active parameters, native tool use and an MXFP4 release intended for local or specialised work.

  • Language & reasoning
  • Coding & agents
Minimum GPU memory
16GB on one GPU
Recommended starting system
Team 32
Licence
Apache License 2.0
See specifications and all 11 systems

Qwen

Qwen3-Coder-Next

An 80-billion-total, 3-billion-active sparse model designed for coding agents, local development and long repository context.

  • Coding & agents
  • Language & reasoning
Minimum GPU memory
160GB aggregate across at least 2 GPUs
Recommended starting system
Frontier Native 2.3TB
Licence
Apache License 2.0
See specifications and all 11 systems

Qwen

Qwen3.5 4B

A compact 4-billion-parameter multimodal model with native vision, reasoning, coding and broad multilingual support.

  • Language & reasoning
  • Coding & agents
  • Vision & OCR
Minimum GPU memory
16GB on one GPU
Recommended starting system
Team 32
Licence
Apache License 2.0
See specifications and all 11 systems

Qwen

Qwen3.5 27B

A 27-billion-parameter dense multimodal model for reasoning, coding, agents and visual understanding across 201 languages and dialects.

  • Language & reasoning
  • Coding & agents
  • Vision & OCR
Minimum GPU memory
64GB on one GPU
Recommended starting system
Studio 96
Licence
Apache License 2.0
See specifications and all 11 systems

Google DeepMind

Gemma 4 12B

Google DeepMind's 12-billion-parameter instruction-tuned Gemma 4 model, supporting text, images, video and audio input with text output.

  • Language & reasoning
  • Vision & OCR
  • Speech & audio
Minimum GPU memory
32GB on one GPU
Recommended starting system
Studio 96
Licence
Apache License 2.0
See specifications and all 11 systems

Google DeepMind

Gemma 4 31B

Google DeepMind's 30.7-billion-parameter instruction-tuned multimodal model for text and image understanding with a 256K context window.

  • Language & reasoning
  • Vision & OCR
Minimum GPU memory
64GB on one GPU
Recommended starting system
Studio 96
Licence
Apache License 2.0
See specifications and all 11 systems

StepFun

Step 3.7 Flash

A 198-billion-parameter sparse vision-language model with about 11 billion active parameters, designed for reasoning, coding, agents and visual intelligence.

  • Language & reasoning
  • Coding & agents
  • Vision & OCR
Minimum GPU memory
416GB aggregate across at least 4 GPUs
Recommended starting system
Specialist configuration required
Licence
Apache License 2.0
See specifications and all 11 systems

Z.ai

GLM-5

A 744-billion-total, 40-billion-active sparse model aimed at complex systems engineering and long-horizon agentic tasks.

  • Language & reasoning
  • Coding & agents
Minimum GPU memory
1536GB aggregate across at least 8 GPUs
Recommended starting system
Specialist configuration required
Licence
MIT License
See specifications and all 11 systems

MiniMax

MiniMax M2.5

A sparse agentic model trained for coding, tool use, search and office work across more than ten programming languages.

  • Language & reasoning
  • Coding & agents
Minimum GPU memory
256GB aggregate across at least 2 GPUs
Recommended starting system
Frontier Native 2.3TB
Licence
Modified MIT License
See specifications and all 11 systems

StepFun

Step-Audio 2 Mini

An end-to-end audio-language model for spoken conversation, audio understanding, paralinguistic cues and audio tool use.

  • Speech & audio
  • Language & reasoning
Minimum GPU memory
24GB on one GPU
Recommended starting system
Team 32
Licence
Apache License 2.0
See specifications and all 11 systems

Model-to-hardware fit

Memory fit is the first gate, not the final recommendation

  1. 01Exact model version

    Use the precise model version, numerical format and complete software pipeline intended for production.

  2. 02Minimum memory

    Check whether it can load on one GPU or requires supported multi-GPU loading.

  3. 03Working headroom

    Allow for context, cache, batch, media encoders and concurrent users.

  4. 04Workload test

    Measure quality, latency and stability on representative work.

What a complete language & reasoning deployment needs

Model weights are only one part of the system. Data access, runtime software, storage, monitoring and administrator handover also affect a reliable deployment.

Diagram showing an approved request, a local service, an approved store and a policy-controlled data path
Diagram showing an approved request, a local service, an approved store and a policy-controlled data path
A private deployment starts with the permitted data path, access policy and logging boundary. GPU Servers technical illustration.
A private deployment starts with the permitted data path, access policy and logging boundary. GPU Servers technical illustration.
GPU server remote management dashboard with system status, access logs and sensor monitoring panels
GPU server remote management dashboard with system status, access logs and sensor monitoring panels
Supplier screenshot of the platform management interface. The final management features and access policy depend on the ordered system. OEM supplier reference image.
Supplier screenshot of the platform management interface. The final management features and access policy depend on the ordered system. OEM supplier reference image.
Diagram showing approved documents moving through a searchable index to an answer with a source citation
Diagram showing approved documents moving through a searchable index to an answer with a source citation
A retrieval workflow should connect each useful answer to approved source material and defined refusal behaviour. GPU Servers technical illustration.
A retrieval workflow should connect each useful answer to approved source material and defined refusal behaviour. GPU Servers technical illustration.
Diagram of an evidence pack containing an asset schedule, burn-in record, health readings, workload test and admin guide
Diagram of an evidence pack containing an asset schedule, burn-in record, health readings, workload test and admin guide
A complete handover includes the supplied assets, test results, operating instructions and agreed follow-up work. GPU Servers technical illustration.
A complete handover includes the supplied assets, test results, operating instructions and agreed follow-up work. GPU Servers technical illustration.