A 2.8-trillion-parameter, 104-billion-active multimodal mixture-of-experts model for long-context reasoning, coding, visual understanding and agentic work.
- Language & reasoning
- Coding & agents
- Vision & OCR
- Minimum GPU memory
- At least 1.561TB aggregate GPU memory for the released weight files
- Recommended starting system
- Specialist configuration required
- Licence
- Kimi K3 Licence
See specifications and all 11 systems → A 36-billion-parameter multimodal mixture-of-experts model with about 3 billion active parameters, a 262,144-token default context and a strong emphasis on agentic coding.
- Language & reasoning
- Coding & agents
- Vision & OCR
- Minimum GPU memory
- 80GB-class GPU for the named BF16 weights at reduced context
- Recommended starting system
- Studio 96
- Licence
- Apache License 2.0
See specifications and all 11 systems → A 284B-total, 13B-active mixture-of-experts language model with one-million-token context and a mixed FP4/FP8 release format.
- Language & reasoning
- Coding & agents
- Minimum GPU memory
- At least 176GB aggregate GPU memory with a supported sharding route
- Recommended starting system
- Frontier Native 2.3TB
- Licence
- MIT License
See specifications and all 11 systems → A 119B-total, 6B-active hybrid model that combines instruction following, reasoning, coding-agent and multimodal capabilities.
- Language & reasoning
- Coding & agents
- Vision & OCR
- Minimum GPU memory
- 256GB aggregate with supported model parallelism
- Recommended starting system
- Frontier Native 2.3TB
- Licence
- Apache License 2.0
See specifications and all 11 systems → Meta's 109B-total, 17B-active multimodal mixture-of-experts checkpoint with text and image input and a very long documented context.
- Language & reasoning
- Vision & OCR
- Minimum GPU memory
- At least 240GB aggregate for BF16 weights and minimal overhead
- Recommended starting system
- Frontier Native 2.3TB
- Licence
- Llama 4 Community Licence
See specifications and all 11 systems → OpenAI's 117B-total, 5.1B-active open-weight reasoning and agentic model, released with native MXFP4 MoE weights.
- Language & reasoning
- Coding & agents
- Minimum GPU memory
- One 80GB GPU
- Recommended starting system
- Studio 96
- Licence
- Apache License 2.0
See specifications and all 11 systems → Google's 27B instruction-tuned multimodal model for text and image input, with a 128K context and broad language coverage.
- Language & reasoning
- Vision & OCR
- Minimum GPU memory
- 64GB GPU for BF16 at a bounded context
- Recommended starting system
- Studio 96
- Licence
- Gemma Terms of Use
See specifications and all 11 systems → A 21-billion-parameter sparse reasoning model with 3.6 billion active parameters, native tool use and an MXFP4 release intended for local or specialised work.
- Language & reasoning
- Coding & agents
- Minimum GPU memory
- 16GB on one GPU
- Recommended starting system
- Team 32
- Licence
- Apache License 2.0
See specifications and all 11 systems → An 80-billion-total, 3-billion-active sparse model designed for coding agents, local development and long repository context.
- Coding & agents
- Language & reasoning
- Minimum GPU memory
- 160GB aggregate across at least 2 GPUs
- Recommended starting system
- Frontier Native 2.3TB
- Licence
- Apache License 2.0
See specifications and all 11 systems → A compact 4-billion-parameter multimodal model with native vision, reasoning, coding and broad multilingual support.
- Language & reasoning
- Coding & agents
- Vision & OCR
- Minimum GPU memory
- 16GB on one GPU
- Recommended starting system
- Team 32
- Licence
- Apache License 2.0
See specifications and all 11 systems → A 27-billion-parameter dense multimodal model for reasoning, coding, agents and visual understanding across 201 languages and dialects.
- Language & reasoning
- Coding & agents
- Vision & OCR
- Minimum GPU memory
- 64GB on one GPU
- Recommended starting system
- Studio 96
- Licence
- Apache License 2.0
See specifications and all 11 systems → Google DeepMind's 12-billion-parameter instruction-tuned Gemma 4 model, supporting text, images, video and audio input with text output.
- Language & reasoning
- Vision & OCR
- Speech & audio
- Minimum GPU memory
- 32GB on one GPU
- Recommended starting system
- Studio 96
- Licence
- Apache License 2.0
See specifications and all 11 systems → Google DeepMind's 30.7-billion-parameter instruction-tuned multimodal model for text and image understanding with a 256K context window.
- Language & reasoning
- Vision & OCR
- Minimum GPU memory
- 64GB on one GPU
- Recommended starting system
- Studio 96
- Licence
- Apache License 2.0
See specifications and all 11 systems → A 198-billion-parameter sparse vision-language model with about 11 billion active parameters, designed for reasoning, coding, agents and visual intelligence.
- Language & reasoning
- Coding & agents
- Vision & OCR
- Minimum GPU memory
- 416GB aggregate across at least 4 GPUs
- Recommended starting system
- Specialist configuration required
- Licence
- Apache License 2.0
See specifications and all 11 systems → A 744-billion-total, 40-billion-active sparse model aimed at complex systems engineering and long-horizon agentic tasks.
- Language & reasoning
- Coding & agents
- Minimum GPU memory
- 1536GB aggregate across at least 8 GPUs
- Recommended starting system
- Specialist configuration required
- Licence
- MIT License
See specifications and all 11 systems → A sparse agentic model trained for coding, tool use, search and office work across more than ten programming languages.
- Language & reasoning
- Coding & agents
- Minimum GPU memory
- 256GB aggregate across at least 2 GPUs
- Recommended starting system
- Frontier Native 2.3TB
- Licence
- Modified MIT License
See specifications and all 11 systems → An end-to-end audio-language model for spoken conversation, audio understanding, paralinguistic cues and audio tool use.
- Speech & audio
- Language & reasoning
- Minimum GPU memory
- 24GB on one GPU
- Recommended starting system
- Team 32
- Licence
- Apache License 2.0
See specifications and all 11 systems →
No model matches that search and workload combination.