Private audio AI

Speech and audio models for local AI

Compare speech recognition, text-to-speech and conversational audio models. Real-time usability depends on the complete audio path, including capture, codec, chunking, inference and playback, not only the checkpoint size.

Selection approach

Choose for the workload, then size the complete deployment

Choose the model class first: transcription, speech generation or end-to-end conversation. Language coverage, diarisation, voice rights and streaming latency can rule out an otherwise capable model.

01

What to compare

  • Transcription, TTS or conversational-audio objective
  • Languages, accents and domain vocabulary
  • Streaming and first-audio latency
  • Voice consent and output-governance needs

02

What changes the hardware

  • Audio chunk size and concurrent streams
  • Codec and feature-extractor overhead
  • CPU audio processing and network jitter
  • GPU headroom for real-time response targets

03

What to test before purchase

  • Word or character error rate on representative audio
  • First-token or first-audio latency
  • Speaker, accent and noisy-room behaviour
  • Stable concurrent streaming over long sessions

6 current models

Compare models by workload and GPU memory

Showing 6 models

Qwen

Qwen3-ASR 1.7B

A compact automatic speech recognition checkpoint released in July 2026 for multilingual transcription and audio understanding.

  • Speech & audio
Minimum GPU memory
8GB GPU planning floor
Recommended starting system
Team 32
Licence
Apache License 2.0
See specifications and all 11 systems

OpenAI

Whisper Large V3

OpenAI's 1.55B-parameter multilingual speech-recognition and translation checkpoint, widely supported across transcription runtimes.

  • Speech & audio
Minimum GPU memory
6GB GPU planning floor for an optimised inference runtime
Recommended starting system
Team 32
Licence
Apache License 2.0
See specifications and all 11 systems

Google DeepMind

Gemma 4 12B

Google DeepMind's 12-billion-parameter instruction-tuned Gemma 4 model, supporting text, images, video and audio input with text output.

  • Language & reasoning
  • Vision & OCR
  • Speech & audio
Minimum GPU memory
32GB on one GPU
Recommended starting system
Studio 96
Licence
Apache License 2.0
See specifications and all 11 systems

Qwen

Qwen3-TTS 1.7B

A multilingual text-to-speech model with streaming and non-streaming generation, instruction-controlled delivery and nine supplied voice timbres.

  • Speech & audio
Minimum GPU memory
8GB on one GPU
Recommended starting system
Team 32
Licence
Apache License 2.0
See specifications and all 11 systems

StepFun

Step-Audio 2 Mini

An end-to-end audio-language model for spoken conversation, audio understanding, paralinguistic cues and audio tool use.

  • Speech & audio
  • Language & reasoning
Minimum GPU memory
24GB on one GPU
Recommended starting system
Team 32
Licence
Apache License 2.0
See specifications and all 11 systems

Model-to-hardware fit

Memory fit is the first gate, not the final recommendation

  1. 01Exact model version

    Use the precise model version, numerical format and complete software pipeline intended for production.

  2. 02Minimum memory

    Check whether it can load on one GPU or requires supported multi-GPU loading.

  3. 03Working headroom

    Allow for context, cache, batch, media encoders and concurrent users.

  4. 04Workload test

    Measure quality, latency and stability on representative work.

What a complete speech & audio deployment needs

Model weights are only one part of the system. Data access, runtime software, storage, monitoring and administrator handover also affect a reliable deployment.

Diagram showing an approved request, a local service, an approved store and a policy-controlled data path
Diagram showing an approved request, a local service, an approved store and a policy-controlled data path
A private deployment starts with the permitted data path, access policy and logging boundary. GPU Servers technical illustration.
A private deployment starts with the permitted data path, access policy and logging boundary. GPU Servers technical illustration.
GPU server remote management dashboard with system status, access logs and sensor monitoring panels
GPU server remote management dashboard with system status, access logs and sensor monitoring panels
Supplier screenshot of the platform management interface. The final management features and access policy depend on the ordered system. OEM supplier reference image.
Supplier screenshot of the platform management interface. The final management features and access policy depend on the ordered system. OEM supplier reference image.
Diagram showing approved documents moving through a searchable index to an answer with a source citation
Diagram showing approved documents moving through a searchable index to an answer with a source citation
A retrieval workflow should connect each useful answer to approved source material and defined refusal behaviour. GPU Servers technical illustration.
A retrieval workflow should connect each useful answer to approved source material and defined refusal behaviour. GPU Servers technical illustration.
Diagram of an evidence pack containing an asset schedule, burn-in record, health readings, workload test and admin guide
Diagram of an evidence pack containing an asset schedule, burn-in record, health readings, workload test and admin guide
A complete handover includes the supplied assets, test results, operating instructions and agreed follow-up work. GPU Servers technical illustration.
A complete handover includes the supplied assets, test results, operating instructions and agreed follow-up work. GPU Servers technical illustration.