A 2.8-trillion-parameter, 104-billion-active multimodal mixture-of-experts model for long-context reasoning, coding, visual understanding and agentic work.
- Language & reasoning
- Coding & agents
- Vision & OCR
- Minimum GPU memory
- At least 1.561TB aggregate GPU memory for the released weight files
- Recommended starting system
- Specialist configuration required
- Licence
- Kimi K3 Licence
See specifications and all 11 systems → A 36-billion-parameter multimodal mixture-of-experts model with about 3 billion active parameters, a 262,144-token default context and a strong emphasis on agentic coding.
- Language & reasoning
- Coding & agents
- Vision & OCR
- Minimum GPU memory
- 80GB-class GPU for the named BF16 weights at reduced context
- Recommended starting system
- Studio 96
- Licence
- Apache License 2.0
See specifications and all 11 systems → A 284B-total, 13B-active mixture-of-experts language model with one-million-token context and a mixed FP4/FP8 release format.
- Language & reasoning
- Coding & agents
- Minimum GPU memory
- At least 176GB aggregate GPU memory with a supported sharding route
- Recommended starting system
- Frontier Native 2.3TB
- Licence
- MIT License
See specifications and all 11 systems → A 119B-total, 6B-active hybrid model that combines instruction following, reasoning, coding-agent and multimodal capabilities.
- Language & reasoning
- Coding & agents
- Vision & OCR
- Minimum GPU memory
- 256GB aggregate with supported model parallelism
- Recommended starting system
- Frontier Native 2.3TB
- Licence
- Apache License 2.0
See specifications and all 11 systems → Meta's 109B-total, 17B-active multimodal mixture-of-experts checkpoint with text and image input and a very long documented context.
- Language & reasoning
- Vision & OCR
- Minimum GPU memory
- At least 240GB aggregate for BF16 weights and minimal overhead
- Recommended starting system
- Frontier Native 2.3TB
- Licence
- Llama 4 Community Licence
See specifications and all 11 systems → OpenAI's 117B-total, 5.1B-active open-weight reasoning and agentic model, released with native MXFP4 MoE weights.
- Language & reasoning
- Coding & agents
- Minimum GPU memory
- One 80GB GPU
- Recommended starting system
- Studio 96
- Licence
- Apache License 2.0
See specifications and all 11 systems → Google's 27B instruction-tuned multimodal model for text and image input, with a 128K context and broad language coverage.
- Language & reasoning
- Vision & OCR
- Minimum GPU memory
- 64GB GPU for BF16 at a bounded context
- Recommended starting system
- Studio 96
- Licence
- Gemma Terms of Use
See specifications and all 11 systems → A 3B-class image-to-text model for document optical character recognition and layout-aware text extraction.
- Minimum GPU memory
- 12GB GPU planning floor
- Recommended starting system
- Team 32
- Licence
- Apache License 2.0
See specifications and all 11 systems → The largest Qwen3 embedding checkpoint, designed for multilingual dense retrieval, classification, clustering and text matching.
- Minimum GPU memory
- 20GB GPU planning floor
- Recommended starting system
- Team 32
- Licence
- Apache License 2.0
See specifications and all 11 systems → A multilingual embedding model that supports dense, sparse and multi-vector retrieval with 1,024 dimensions and inputs up to 8,192 tokens.
- Minimum GPU memory
- 4GB GPU planning floor, or CPU for light use
- Recommended starting system
- Team 32
- Licence
- MIT License
See specifications and all 11 systems → A compact automatic speech recognition checkpoint released in July 2026 for multilingual transcription and audio understanding.
- Minimum GPU memory
- 8GB GPU planning floor
- Recommended starting system
- Team 32
- Licence
- Apache License 2.0
See specifications and all 11 systems → OpenAI's 1.55B-parameter multilingual speech-recognition and translation checkpoint, widely supported across transcription runtimes.
- Minimum GPU memory
- 6GB GPU planning floor for an optimised inference runtime
- Recommended starting system
- Team 32
- Licence
- Apache License 2.0
See specifications and all 11 systems → A 4B-class real-time automatic speech recognition model released by Mistral AI for low-latency streaming transcription.
- Minimum GPU memory
- 24GB GPU planning floor
- Recommended starting system
- Team 32
- Licence
- Apache License 2.0
See specifications and all 11 systems → A compact rectified-flow model for text-to-image, image editing and multi-reference work, released under Apache 2.0.
- Minimum GPU memory
- About 13GB VRAM
- Recommended starting system
- Team 32
- Licence
- Apache License 2.0
See specifications and all 11 systems → The December 2025 Qwen-Image update for text-to-image generation, with improved realism, natural detail and text rendering.
- Minimum GPU memory
- 64GB GPU as a tight full-weight floor
- Recommended starting system
- Studio 96
- Licence
- Apache License 2.0
See specifications and all 11 systems → A 5B text-and-image-to-video model that supports 720p generation and an official single-GPU offload route.
- Minimum GPU memory
- 24GB VRAM with documented CPU offload settings
- Recommended starting system
- Studio 96
- Licence
- Apache License 2.0
See specifications and all 11 systems → An 8.3B text-to-video and image-to-video model with 480p and 720p checkpoints, optional super-resolution and distilled workflows.
- Minimum GPU memory
- 14GB VRAM with model offloading
- Recommended starting system
- Team 32
- Licence
- Tencent Hunyuan Community Licence
See specifications and all 11 systems → Tencent's image-to-3D generation pipeline for shape and texture creation, with open model weights and a model-specific community licence.
- Minimum GPU memory
- 24GB single-GPU planning floor
- Recommended starting system
- Studio 96
- Licence
- Tencent Hunyuan Community Licence
See specifications and all 11 systems → A 12B multimodal safeguard model for classifying text and image prompts and responses against Meta's hazard taxonomy.
- Safety & moderation
- Vision & OCR
- Minimum GPU memory
- 28GB GPU planning floor
- Recommended starting system
- Team 32
- Licence
- Llama 4 Community Licence
See specifications and all 11 systems → An 8B-class text-to-image model in the Stable Diffusion 3.5 family, with a large ecosystem of Diffusers and ComfyUI workflows.
- Minimum GPU memory
- 24GB planning floor with an optimised or offload workflow
- Recommended starting system
- Team 32
- Licence
- Stability AI Community Licence
See specifications and all 11 systems → A 21-billion-parameter sparse reasoning model with 3.6 billion active parameters, native tool use and an MXFP4 release intended for local or specialised work.
- Language & reasoning
- Coding & agents
- Minimum GPU memory
- 16GB on one GPU
- Recommended starting system
- Team 32
- Licence
- Apache License 2.0
See specifications and all 11 systems → An 80-billion-total, 3-billion-active sparse model designed for coding agents, local development and long repository context.
- Coding & agents
- Language & reasoning
- Minimum GPU memory
- 160GB aggregate across at least 2 GPUs
- Recommended starting system
- Frontier Native 2.3TB
- Licence
- Apache License 2.0
See specifications and all 11 systems → A compact 4-billion-parameter multimodal model with native vision, reasoning, coding and broad multilingual support.
- Language & reasoning
- Coding & agents
- Vision & OCR
- Minimum GPU memory
- 16GB on one GPU
- Recommended starting system
- Team 32
- Licence
- Apache License 2.0
See specifications and all 11 systems → A 27-billion-parameter dense multimodal model for reasoning, coding, agents and visual understanding across 201 languages and dialects.
- Language & reasoning
- Coding & agents
- Vision & OCR
- Minimum GPU memory
- 64GB on one GPU
- Recommended starting system
- Studio 96
- Licence
- Apache License 2.0
See specifications and all 11 systems → Google DeepMind's 12-billion-parameter instruction-tuned Gemma 4 model, supporting text, images, video and audio input with text output.
- Language & reasoning
- Vision & OCR
- Speech & audio
- Minimum GPU memory
- 32GB on one GPU
- Recommended starting system
- Studio 96
- Licence
- Apache License 2.0
See specifications and all 11 systems → Google DeepMind's 30.7-billion-parameter instruction-tuned multimodal model for text and image understanding with a 256K context window.
- Language & reasoning
- Vision & OCR
- Minimum GPU memory
- 64GB on one GPU
- Recommended starting system
- Studio 96
- Licence
- Apache License 2.0
See specifications and all 11 systems → A 198-billion-parameter sparse vision-language model with about 11 billion active parameters, designed for reasoning, coding, agents and visual intelligence.
- Language & reasoning
- Coding & agents
- Vision & OCR
- Minimum GPU memory
- 416GB aggregate across at least 4 GPUs
- Recommended starting system
- Specialist configuration required
- Licence
- Apache License 2.0
See specifications and all 11 systems → A 744-billion-total, 40-billion-active sparse model aimed at complex systems engineering and long-horizon agentic tasks.
- Language & reasoning
- Coding & agents
- Minimum GPU memory
- 1536GB aggregate across at least 8 GPUs
- Recommended starting system
- Specialist configuration required
- Licence
- MIT License
See specifications and all 11 systems → A sparse agentic model trained for coding, tool use, search and office work across more than ten programming languages.
- Language & reasoning
- Coding & agents
- Minimum GPU memory
- 256GB aggregate across at least 2 GPUs
- Recommended starting system
- Frontier Native 2.3TB
- Licence
- Modified MIT License
See specifications and all 11 systems → A multilingual text-to-speech model with streaming and non-streaming generation, instruction-controlled delivery and nine supplied voice timbres.
- Minimum GPU memory
- 8GB on one GPU
- Recommended starting system
- Team 32
- Licence
- Apache License 2.0
See specifications and all 11 systems → An 8-billion-parameter multilingual reranker for scoring retrieved passages across more than 100 natural and programming languages.
- Minimum GPU memory
- 24GB on one GPU
- Recommended starting system
- Team 32
- Licence
- Apache License 2.0
See specifications and all 11 systems → A 9-billion-parameter rectified-flow image model for text-to-image generation and multi-reference image editing.
- Minimum GPU memory
- 64GB on one GPU
- Recommended starting system
- Studio 96
- Licence
- FLUX Non-Commercial License
See specifications and all 11 systems → An end-to-end audio-language model for spoken conversation, audio understanding, paralinguistic cues and audio tool use.
- Speech & audio
- Language & reasoning
- Minimum GPU memory
- 24GB on one GPU
- Recommended starting system
- Team 32
- Licence
- Apache License 2.0
See specifications and all 11 systems →
No model matches that search and workload combination.