Developer AI

Coding and agent models for local deployment

Compare models for code generation, repository reasoning, terminal work and tool-driven agents. Long context helps, but reliable tool use, edit quality and runtime integration usually matter more than a headline benchmark.

Selection approach

Choose for the workload, then size the complete deployment

Select against the repositories, languages, test suites and agent framework you will actually use. A model that fits memory can still be a poor choice if tool calls, patch quality or long-running task recovery are weak.

01

What to compare

  • Repository and programming-language coverage
  • Tool-call format and agent-framework compatibility
  • Patch accuracy, test repair and instruction following
  • Context management for large codebases

02

What changes the hardware

  • Prompt-prefill cost for long repositories
  • Active versus total parameters in sparse models
  • Runtime support for tool-call and reasoning formats
  • Replica capacity for several developers

03

What to test before purchase

  • Resolve representative issues without leaking test answers
  • Produce patches that compile and pass the existing suite
  • Recover from failed tools and interrupted sessions
  • Measure latency at realistic repository context

12 current models

Compare models by workload and GPU memory

Showing 12 models

Moonshot AI

Kimi K3

A 2.8-trillion-parameter, 104-billion-active multimodal mixture-of-experts model for long-context reasoning, coding, visual understanding and agentic work.

  • Language & reasoning
  • Coding & agents
  • Vision & OCR
Minimum GPU memory
At least 1.561TB aggregate GPU memory for the released weight files
Recommended starting system
Specialist configuration required
Licence
Kimi K3 Licence
See specifications and all 11 systems

Qwen

Qwen3.6 35B-A3B

A 36-billion-parameter multimodal mixture-of-experts model with about 3 billion active parameters, a 262,144-token default context and a strong emphasis on agentic coding.

  • Language & reasoning
  • Coding & agents
  • Vision & OCR
Minimum GPU memory
80GB-class GPU for the named BF16 weights at reduced context
Recommended starting system
Studio 96
Licence
Apache License 2.0
See specifications and all 11 systems

DeepSeek

DeepSeek V4 Flash

A 284B-total, 13B-active mixture-of-experts language model with one-million-token context and a mixed FP4/FP8 release format.

  • Language & reasoning
  • Coding & agents
Minimum GPU memory
At least 176GB aggregate GPU memory with a supported sharding route
Recommended starting system
Frontier Native 2.3TB
Licence
MIT License
See specifications and all 11 systems

Mistral AI

Mistral Small 4

A 119B-total, 6B-active hybrid model that combines instruction following, reasoning, coding-agent and multimodal capabilities.

  • Language & reasoning
  • Coding & agents
  • Vision & OCR
Minimum GPU memory
256GB aggregate with supported model parallelism
Recommended starting system
Frontier Native 2.3TB
Licence
Apache License 2.0
See specifications and all 11 systems

OpenAI

GPT-OSS 120B

OpenAI's 117B-total, 5.1B-active open-weight reasoning and agentic model, released with native MXFP4 MoE weights.

  • Language & reasoning
  • Coding & agents
Minimum GPU memory
One 80GB GPU
Recommended starting system
Studio 96
Licence
Apache License 2.0
See specifications and all 11 systems

OpenAI

GPT-OSS 20B

A 21-billion-parameter sparse reasoning model with 3.6 billion active parameters, native tool use and an MXFP4 release intended for local or specialised work.

  • Language & reasoning
  • Coding & agents
Minimum GPU memory
16GB on one GPU
Recommended starting system
Team 32
Licence
Apache License 2.0
See specifications and all 11 systems

Qwen

Qwen3-Coder-Next

An 80-billion-total, 3-billion-active sparse model designed for coding agents, local development and long repository context.

  • Coding & agents
  • Language & reasoning
Minimum GPU memory
160GB aggregate across at least 2 GPUs
Recommended starting system
Frontier Native 2.3TB
Licence
Apache License 2.0
See specifications and all 11 systems

Qwen

Qwen3.5 4B

A compact 4-billion-parameter multimodal model with native vision, reasoning, coding and broad multilingual support.

  • Language & reasoning
  • Coding & agents
  • Vision & OCR
Minimum GPU memory
16GB on one GPU
Recommended starting system
Team 32
Licence
Apache License 2.0
See specifications and all 11 systems

Qwen

Qwen3.5 27B

A 27-billion-parameter dense multimodal model for reasoning, coding, agents and visual understanding across 201 languages and dialects.

  • Language & reasoning
  • Coding & agents
  • Vision & OCR
Minimum GPU memory
64GB on one GPU
Recommended starting system
Studio 96
Licence
Apache License 2.0
See specifications and all 11 systems

StepFun

Step 3.7 Flash

A 198-billion-parameter sparse vision-language model with about 11 billion active parameters, designed for reasoning, coding, agents and visual intelligence.

  • Language & reasoning
  • Coding & agents
  • Vision & OCR
Minimum GPU memory
416GB aggregate across at least 4 GPUs
Recommended starting system
Specialist configuration required
Licence
Apache License 2.0
See specifications and all 11 systems

Z.ai

GLM-5

A 744-billion-total, 40-billion-active sparse model aimed at complex systems engineering and long-horizon agentic tasks.

  • Language & reasoning
  • Coding & agents
Minimum GPU memory
1536GB aggregate across at least 8 GPUs
Recommended starting system
Specialist configuration required
Licence
MIT License
See specifications and all 11 systems

MiniMax

MiniMax M2.5

A sparse agentic model trained for coding, tool use, search and office work across more than ten programming languages.

  • Language & reasoning
  • Coding & agents
Minimum GPU memory
256GB aggregate across at least 2 GPUs
Recommended starting system
Frontier Native 2.3TB
Licence
Modified MIT License
See specifications and all 11 systems

Model-to-hardware fit

Memory fit is the first gate, not the final recommendation

  1. 01Exact model version

    Use the precise model version, numerical format and complete software pipeline intended for production.

  2. 02Minimum memory

    Check whether it can load on one GPU or requires supported multi-GPU loading.

  3. 03Working headroom

    Allow for context, cache, batch, media encoders and concurrent users.

  4. 04Workload test

    Measure quality, latency and stability on representative work.

What a complete coding & agents deployment needs

Model weights are only one part of the system. Data access, runtime software, storage, monitoring and administrator handover also affect a reliable deployment.

Diagram showing an approved request, a local service, an approved store and a policy-controlled data path
Diagram showing an approved request, a local service, an approved store and a policy-controlled data path
A private deployment starts with the permitted data path, access policy and logging boundary. GPU Servers technical illustration.
A private deployment starts with the permitted data path, access policy and logging boundary. GPU Servers technical illustration.
GPU server remote management dashboard with system status, access logs and sensor monitoring panels
GPU server remote management dashboard with system status, access logs and sensor monitoring panels
Supplier screenshot of the platform management interface. The final management features and access policy depend on the ordered system. OEM supplier reference image.
Supplier screenshot of the platform management interface. The final management features and access policy depend on the ordered system. OEM supplier reference image.
Diagram showing approved documents moving through a searchable index to an answer with a source citation
Diagram showing approved documents moving through a searchable index to an answer with a source citation
A retrieval workflow should connect each useful answer to approved source material and defined refusal behaviour. GPU Servers technical illustration.
A retrieval workflow should connect each useful answer to approved source material and defined refusal behaviour. GPU Servers technical illustration.
Diagram of an evidence pack containing an asset schedule, burn-in record, health readings, workload test and admin guide
Diagram of an evidence pack containing an asset schedule, burn-in record, health readings, workload test and admin guide
A complete handover includes the supplied assets, test results, operating instructions and agreed follow-up work. GPU Servers technical illustration.
A complete handover includes the supplied assets, test results, operating instructions and agreed follow-up work. GPU Servers technical illustration.