Best larger-system route
Team 32
The smallest system in the range that reaches this model's preferred working allowance.
Explore Team 32
Qwen model
The BF16 files total about 15.1GB. One 32GB GPU is a comfortable starting point for a single embedding worker; product sizing should then follow tokens per second, document backlog and concurrent indexing rather than model fit alone.
Qwen/Qwen3-Embedding-8B1d8ad4ca9b3dQwen/Qwen3-Embedding-8B Buyer verdict
Use Qwen3-Embedding 8B when multilingual retrieval quality and long input support matter more than the lower memory and latency of the 0.6B or 4B variants.
These figures apply to the named model version. Quantisation, fine-tuning, context length, image resolution, batch size and serving software can materially change the hardware needed.
Hardware requirements
The floor adds limited runtime and batch headroom above the 15.1GB weight files.
Batch, sequence length and the service-level target should still be measured.
Technical specification
Specifications shown for source version
1d8ad4ca9b3dd8059ad90a75d4983776a23d44af, updated
7 July 2025.
Product compatibility
A system is listed as fitting when its GPU memory meets the requirement shown above. Speed, usable context, batch size and concurrent users still need testing with the final model and software configuration.
Best larger-system route
The smallest system in the range that reaches this model's preferred working allowance.
Explore Team 32sme workstation
Also meets this model page's recommended working allowance.
Available GPU memory exceeds the calculated or estimated requirement. Confirm the final precision, software, workload and performance before purchase.
sme workstation
Also meets this model page's recommended working allowance.
Available GPU memory exceeds the calculated or estimated requirement. Confirm the final precision, software, workload and performance before purchase.
sme workstation
Also meets this model page's recommended working allowance.
Available GPU memory exceeds the calculated or estimated requirement. Confirm the final precision, software, workload and performance before purchase.
sme workstation
Also meets this model page's recommended working allowance.
Available GPU memory exceeds the calculated or estimated requirement. Confirm the final precision, software, workload and performance before purchase.
pcie rack
Also meets this model page's recommended working allowance.
Available GPU memory exceeds the calculated or estimated requirement. Confirm the final precision, software, workload and performance before purchase.
pcie rack
Also meets this model page's recommended working allowance.
Available GPU memory exceeds the calculated or estimated requirement. Confirm the final precision, software, workload and performance before purchase.
pcie rack
Also meets this model page's recommended working allowance.
Available GPU memory exceeds the calculated or estimated requirement. Confirm the final precision, software, workload and performance before purchase.
pcie rack
Also meets this model page's recommended working allowance.
Available GPU memory exceeds the calculated or estimated requirement. Confirm the final precision, software, workload and performance before purchase.
pcie rack
Also meets this model page's recommended working allowance.
Available GPU memory exceeds the calculated or estimated requirement. Confirm the final precision, software, workload and performance before purchase.
frontier partner
Also meets this model page's recommended working allowance.
Available GPU memory exceeds the calculated or estimated requirement. Confirm the final precision, software, workload and performance before purchase.
frontier partner
Also meets this model page's recommended working allowance.
Available GPU memory exceeds the calculated or estimated requirement. Confirm the final precision, software, workload and performance before purchase.
Deployment reality
Serving software
Before installation
System requirements
Performance: speed, usable context and concurrency depend on the selected system, software and workload. No benchmark is quoted on this page.
Commercial and legal boundary
Commercial use: permitted
Apache 2.0 permits commercial use subject to notices.
Always retain the applicable notices and recheck the live terms for the intended organisation, territory, use and distribution route. Obtain legal advice where required; the official licence governs use.
Read the official licenceTechnical questions
Its memory requirement fits. Actual batch, token length and throughput still need testing.
No. Smaller series variants may provide adequate retrieval quality with lower latency and more room for other services.
No. Chunking should be designed around retrieval relevance and document structure, not the model ceiling.
GPU memory is only one part of the system. Storage, data access, serving software, monitoring and administrator handover also affect a reliable deployment.
Privacy & cookies. Google Analytics measures site use so we can improve it. Read our Cookie Notice and Privacy Notice. To object, disable cookies in your browser.