Choose an AI server by proving a business workload on a controlled configuration. Specification sheets create a shortlist; they do not supply the acceptance result.
Write the acceptance set first
Collect representative, authorised inputs and expected outcomes. Include ordinary work, difficult cases, refusal cases and security boundaries.
Name the decision metrics. For retrieval this may be supported-answer rate and permission preservation. For transcription it may be word error rate and real-time factor. For interactive inference it may include first-token and tail latency.
Freeze the test identity
Record:
- product serial or reference BOM;
- GPU model, edition, count and topology;
- CPU, RAM, storage and network;
- firmware, operating system and power mode;
- driver and CUDA versions;
- runtime, container digest and launch settings;
- exact checkpoint, revision and quantisation;
- input, output, context, batch and concurrency.
Without this identity, a result cannot be reproduced or attached safely to a product.
Separate quality from load
First confirm that the model and application pass the task set. Then run load tests using the same service path.
A faster model that fails the business task is not a better result. A high-quality model that misses the latency or capacity gate may still require a different deployment.
Use a load curve
Test several arrival rates and concurrency levels. Record completed requests, errors, time-to-first-token, end-to-end latency, output rate, memory, power and temperature.
Look for the knee where queueing and tail latency begin to deteriorate. Design the operating limit below failure, with room for variation and maintenance.
Measure the entire path
Browser, API gateway, retrieval, storage and network can dominate user experience. Run both engine-level benchmarks and application-level transactions.
For multi-GPU systems, capture topology and validate the intended parallelism. For a rack system, include facility and network conditions that can change throttling or availability.
Test stability and recovery
Run for a representative duration. Exercise restart, failed request, full queue, monitoring alert and rollback. Confirm what happens after a driver, runtime or model update.
Supportability includes logs, spares, warranty route and a named operator, not just uptime during the sales demonstration.
Compare with hosted and smaller options
Run the accepted workload through a suitable hosted service where policy permits. Price the real usage and integration. Also test the next smaller local configuration.
The correct result may be hosted, hybrid or a smaller server. A benchmark process that can only recommend the largest system is not a decision process.
Evidence states
Keep conclusions explicit:
- measured support: exact configuration passed the recorded gate;
- vendor documented: a primary source states a specification;
- calculated candidate: arithmetic clears a defined capacity check;
- assumption: plausible but not validated;
- not evaluated: no controlled result exists.
Publish the conditions beside the result and set a review date. Model repositories, runtimes, suppliers and prices change.
Purchase gate
Approve a system only when workload quality, service performance, facility fit, operations, support and economics are acceptable together. Retain the raw logs and the configuration manifest as part of the product record.
Technical context
See the physical and operating boundary
Use these views to connect the guide to the machine, its airflow and its operating environment. Captions state the limits of what each image shows.
Primary sources
Sources are checked at the review date. Platform terms, prices and public guidance can change; verify them at the point of decision.