Workload notebook

Start with a bounded workload, not a server model.

The same GPU count can serve three very different jobs. Define the input, demand pattern and acceptance test first, then decide whether owned capacity is proportionate.

Diagram showing an approved request, a local service, an approved store and a policy-controlled data path
Diagram showing an approved request, a local service, an approved store and a policy-controlled data path
A private deployment starts with the permitted data path, access policy and logging boundary. GPU Servers technical illustration.
A private deployment starts with the permitted data path, access policy and logging boundary.

Before hardware

Five questions that change the design

If these answers are missing, a larger specification only makes the uncertainty more expensive.

  1. What data enters?

    Name the document, repository, media or batch source. Record its owner, sensitivity, size and update pattern.

  2. What must the output pass?

    Build a representative acceptance set before discussing a package. A plausible demo is not a pass condition.

  3. How is demand shaped?

    Concurrent users, context, queue length, latency and working hours drive different capacity decisions.

  4. Who operates it?

    Name the owner for access, updates, logs, backup, incidents and the hosted fallback.

  5. Where will it run?

    Check rack space, circuits, heat, noise, network and physical access before treating the design as viable.

Workload patterns

Choose by the work your team needs to finish

Each pattern identifies the inputs, demand, acceptance test and hosted alternative that should be considered before hardware is selected.

Workload 01

private RAG server

A private document assistant that shows its sources and its limits.

Build a private company document assistant with source citations, permission-aware retrieval, evaluation and a customer-controlled server.

Bring to the sizing session
Approved documents, permission groups, question set and update frequency
Acceptance evidence
Retrieval, citation support, refusal and permission boundaries
Hosted route
Choose hosted when connector maturity or managed model quality matters more than a local route.
Examine this workload
Workload 02

local AI coding server

Local code assistance measured against your repositories.

Private AI coding servers for repository-aware assistance, controlled source-code access, model evaluation and multi-developer use.

Bring to the sizing session
Repository scope, developer concurrency, context length and secret boundaries
Acceptance evidence
Correctness, security, review effort and useful response time
Hosted route
Keep hosted coding tools where their capability and commercial terms fit the repositories.
Examine this workload
Workload 03

GPU rendering server

Dense GPU workers for jobs that do not need one shared memory pool.

Physical GPU servers for rendering, image, transcription, embeddings and independent batch workers - with power and marketplace cautions.

Bring to the sizing session
Job queue, software licences, completion target, wall power and working hours
Acceptance evidence
Completed useful work, queue time, stability and measured energy
Hosted route
Use rented capacity where work is irregular or deadlines need rapid scale.
Examine this workload
Workload 04

private OCR server

Document extraction measured field by field.

Evaluate private OCR, document extraction and classification with a versioned document set, field-level accuracy and controlled source retention.

Bring to the sizing session
Authorised documents, layouts, languages, labelled fields and exception rules
Acceptance evidence
Field or character accuracy, latency, provenance and exception handling
Hosted route
Choose hosted document AI where connector maturity and elastic page volume outweigh a local data route.
Examine this workload
Workload 05

private transcription server

Transcription judged by the words it gets wrong.

Plan local transcription with word error rate, real-time factor, speaker handling, privacy and a representative audio evaluation set.

Bring to the sizing session
Authorised audio, languages, accents, channels, noise and live or batch target
Acceptance evidence
Word error rate, real-time factor, latency, failures and correction effort
Hosted route
Use hosted transcription for irregular demand where provider terms and the audio policy are acceptable.
Examine this workload
Workload 06

private embedding server

Embedding capacity tied to retrieval quality.

Design private embedding and retrieval capacity around exact checkpoints, dimensions, corpus updates, query latency and retrieval-quality tests.

Bring to the sizing session
Exact checkpoint, corpus, update frequency, dimensions, queries and known relevant passages
Acceptance evidence
Retrieval quality, vectors per second, query latency and index rebuild time
Hosted route
Use a managed vector or embedding service where its controls, integrations and variable demand fit.
Examine this workload
Workload 07

private AI image generation server

Image generation measured by accepted outputs.

Plan local image generation around exact checkpoints, licences, resolution, steps, accepted-output rate, storage and creative review.

Bring to the sizing session
Exact checkpoint, rights, prompts, resolution, steps, sampler, adapters and output review
Acceptance evidence
Accepted-output rate, latency, throughput, stability, memory and storage
Hosted route
Keep cloud burst for irregular campaigns or approved models not validated on the local system.
Examine this workload
Workload 08

private AI video generation server

Video capacity tied to a controlled clip recipe.

Evaluate local AI video generation by exact model, dimensions, frames, duration, stability, accepted clips, storage and facility demand.

Bring to the sizing session
Exact checkpoint, clip dimensions, frames, duration, steps, queue and storage
Acceptance evidence
Accepted clips, clip latency, clips per hour, failures, memory, power and temperature
Hosted route
Rented burst compute may be proportionate for campaign peaks and long jobs.
Examine this workload
Workload 09

private AI fine tuning server

Fine-tuning sized by method, data and evaluation.

Plan controlled fine-tuning around the exact base checkpoint, method, dataset, sequence length, memory, training stability and evaluation.

Bring to the sizing session
Base checkpoint, licence, method, dataset, sequence length, batch and evaluation split
Acceptance evidence
Completed steps, memory, stability, loss, held-out quality and regression checks
Hosted route
Use rented training capacity when experiments are large, infrequent or need rapid cluster scale.
Examine this workload
Workload 10

private AI agents server

Agents bounded by tools, permissions and evidence.

Design private agentic workflows with bounded tools, permissions, audit logs, task-success tests, cost controls and human approval.

Bring to the sizing session
Task set, tools, identities, permitted actions, concurrency, approval and stop conditions
Acceptance evidence
Task success, unsafe actions, retries, tool errors, latency and reviewer effort
Hosted route
Retain approved hosted models for complex exceptions where their capability justifies the data route.
Examine this workload

Evidence sequence

A package follows the acceptance set

This order prevents a polished hardware specification from becoming a substitute for workload proof.

  1. 01

    Collect representative inputs

    Use authorised examples that reflect difficult and ordinary work.

  2. 02

    Agree pass and refusal conditions

    Name the quality, latency, security and human-review requirements.

  3. 03

    Test a reference configuration

    Record model, runtime, context, concurrency, output and wall power.

  4. 04

    Confirm the operating site

    Validate facilities, support ownership, network and fallback route.

Workload decisions become physical

GPU layout, cooling, remote management and the installation setting all affect whether a reference workload can become a dependable business service.

Three-quarter supplier render of a 4U OEM multi-GPU rack server
Three-quarter supplier render of a 4U OEM multi-GPU rack server
OEM platform reference render. It is not evidence of a completed customer build or final specification. OEM supplier reference image.
OEM platform reference render. It is not evidence of a completed customer build or final specification. OEM supplier reference image.
Open 4U OEM GPU server chassis showing passive GPUs, cooling fans, processors and memory slots
Open 4U OEM GPU server chassis showing passive GPUs, cooling fans, processors and memory slots
OEM supplier render showing one possible internal layout. Components vary with the ordered build. OEM supplier reference image.
OEM supplier render showing one possible internal layout. Components vary with the ordered build. OEM supplier reference image.
Diagram combining model weights, context, cache and active requests into a memory headroom check
Diagram combining model weights, context, cache and active requests into a memory headroom check
Memory fit depends on the workload, context, cache and simultaneous demand, then needs a representative test. GPU Servers technical illustration.
Memory fit depends on the workload, context, cache and simultaneous demand, then needs a representative test. GPU Servers technical illustration.
Exploded supplier render of passive GPUs arranged above an open 4U rack chassis
Exploded supplier render of passive GPUs arranged above an open 4U rack chassis
Supplier layout render used to explain GPU density and airflow. It does not represent a confirmed package configuration. OEM supplier reference image.
Supplier layout render used to explain GPU density and airflow. It does not represent a confirmed package configuration. OEM supplier reference image.

Smallest useful next step

Bring one workload, one owner and one pass condition

Do not upload confidential material through the public site. Start by describing the workload without sensitive examples, then agree how protected material will be supplied and tested separately.