Not every team needs a rack, three-phase power and a server room. The task is often simpler than that: keep the model close, work with it every day and let the data stay where it is. That is what a deskside tower with a single 96 GB card is for.
Why 96 GB on one card matters
A model that fits into the memory of a single card is easy to live with: you load it and you run it. The moment it stops fitting, a different kind of engineering begins: splitting layers across cards, passing intermediate results between them, tuning that eats up days. In 96 GB you can comfortably keep models of up to 70 billion parameters in four-bit quantisation, while models up to 32B fit in more precise formats.
There is a second difference that people notice later. The card in a deskside station is the same RTX PRO 6000 Blackwell that goes into servers, only in a workstation build with its own cooling. It is not a cut-down version: the memory and the compute blocks are identical.
Configuration
The next step after DGX Spark
Plenty of teams start with a DGX Spark: a box holding 128 GB of unified memory that sits on the desk and lets you try large models without a server room. Its limit is a single well-known one: memory bandwidth of 273 GB/s. A 70B model in a dense format returns a few tokens per second, and that is too slow to work with.
A station built on the RTX PRO 6000 lifts exactly that limit: Blackwell carries GDDR7 memory with several times the bandwidth. A model of the same size starts answering at a speed you can use daily, rather than one you can only demonstrate.
The Spark stays useful: it is quiet, compact and holds the role of an experiment box well. You need the station once the model has stopped being a toy and become a tool.
What Threadripper PRO gives you over an ordinary processor
Three things, and every one of them is critical for this particular job. Memory with error correction: a long computation should not fall over because of one flipped cell. Eight memory channels instead of two: data reaches the card faster. And a generous number of PCIe lanes, so a second card, a fast drive and the network never take lanes away from each other.
The thirty-two cores are here for data preparation, not for rendering. Tokenisation, loading the dataset, pre-processing images: all of that is processor work, and while it runs the card sits idle.
Room to grow
- A second RTX PRO 6000 Max-Q card: 192 GB of video memory in total. The Max-Q version draws less power and shares a tower far more peacefully.
- Memory up to 512 GB, for the point where datasets stop fitting.
- The power supply was chosen with headroom from the start, so an upgrade needs no new case and no new PSU.
Who it suits and who it does not
A good fit
- Running models of up to 70B locally, when the data must not leave the building.
- Fine-tuning for your own subject area on small datasets.
- Computer vision: training detectors, processing video, generating synthetic data.
- Teams with no server room: the station runs off an ordinary socket and stands in the office.
Better to look elsewhere
- Several people working at once: one card quickly turns into a queue, and a server that splits the GPU through MIG wins here.
- Training from scratch on a large corpus: that calls for shared memory and several accelerators.
Common questions
Is a station like this noisy in an office?
Under load you hear it, but the level is that of a powerful desktop, not of a rack. Server machines have no place in an office at all: they use front-to-back airflow and are in a different noise class entirely.
Is an ordinary socket enough?
Yes, with one card. If you are planning a second one, check the circuit first: peak draw climbs towards one and a half kilowatts.
Why ECC memory when it costs more?
Because the computation runs for hours. A single flipped bit spoils the result quietly, without crashing anything, and you find out about it at the very end.
Can I fit two standard RTX PRO 6000 cards instead of Max-Q?
Technically yes, in practice Max-Q is the better choice: it draws less power and puts out less heat, whereas two full-power cards in a tower sit flush against each other and block one another from breathing.
What we have of this in stock
- ETE AI Workstation: the station described in this article
- ETE AI Server RTX PRO: the same card, four of them, in a rack
- ETE AI Server H200: for when the task has outgrown a single machine
- NVIDIA graphics cards: the RTX PRO 6000 and the Max-Q version on their own