One card gives you 48 GB and runs on 300 W, the other gives 96 GB and asks for 600. Between them sit a single generation, twice the memory, a new FP4 precision and a board that will not drop into every chassis. Here is what the surcharge actually buys, and where RTX 6000 Ada still closes the requirement.

In brief

  • 96 GB of GDDR7 against 48 GB of GDDR6, bandwidth of 1792 against 960 GB/s
  • 24,064 CUDA cores against 18,176, FP32 of 126 against 91.1 TFLOPS
  • Fifth-generation Tensor Cores handle FP4. On Ada the ceiling is FP8
  • 600 W against 300 W, plus a board an inch taller and an inch and a half longer
  • MIG splits Blackwell into four 24 GB instances. Ada has no hardware partitioning at all
  • Llama 3.3 70B in Q4 runs on a single card: 27 tokens per second

What changed in one generation

Spec RTX 6000 Ada RTX PRO 6000 Blackwell WS
Memory 48 GB GDDR6 with ECC 96 GB GDDR7 with ECC
Memory bus 384-bit 512-bit
Bandwidth 960 GB/s 1792 GB/s
CUDA cores 18,176 24,064
Tensor / RT cores 4th and 3rd generation 5th and 4th generation
FP32 91.1 TFLOPS 126 TFLOPS
RT Core 210.6 TFLOPS 382 TFLOPS
Headline AI figure 1457 effective FP8 TFLOPS with sparsity 4000 AI TOPS in FP4 with sparsity
System interface PCIe 4.0 ×16 PCIe 5.0 ×16
Displays 4× DisplayPort 1.4a 4× DisplayPort 2.1b
Video engines 3× NVENC, 3× NVDEC 4× 9th-gen NVENC, 4× 6th-gen NVDEC
MIG none up to 4×24 GB, 2×48 GB or 1×96 GB
TDP 300 W 600 W
Size, slots 4.4 × 10.5 inches, 2 slots 5.4 × 12 inches, 2 slots
CUDA / Vulkan 11.6 / 1.3 12.8 / 1.4

Sources: RTX 6000 Ada datasheet, RTX PRO 6000 Blackwell Workstation Edition datasheet

FP32 grew by 38%, RT Core by 81%, memory bandwidth by 87%. That last number is the one to plan around.

Memory counts for more than cores

Into 48 GB a 70-billion-parameter model fits only under aggressive quantization and with a short context, and the KV cache has to be trimmed. Into 96 GB the same model drops in at Q4 with a working context, and there is room left over. llama-bench figures on a single request, RTX PRO 6000:

Model Prompt reading Generation
llama3.1 70B Q4_K_M not reported 27 tok/s
qwen3-next 80B Q4_K_M 3274 tok/s 124 tok/s
qwen2.5 32B not reported 56 tok/s
mistral-nemo 12B 8325 tok/s 158 tok/s

Source: llama-bench on RTX PRO 6000, single-user mode

There is a trap here that budget planning walks into regularly. Two RTX 6000 Ada boards do not add up to 96 GB: the datasheet prints a flat “No” next to NVLink, traffic goes over PCIe, and a model that missed the fit on one card has to be split across layers at a speed penalty. Practitioners frame the second card differently: keep several specialized models resident at once, say an LLM, a reranker and an embedding model.

FP4 and fifth-generation Tensor Cores

FP4 halves the memory footprint of FP8, and that is exactly where the headline 4000 AI TOPS comes from. The datasheet footnote reads “effective FP4 TOPS with sparsity”, so on a dense model without sparsity the figure lands lower. The gain from lower precision also arrives from an unexpected direction: token generation is bound by memory speed, because every token requires reading all the weights. The precision that wins is the one that makes the model physically smaller.

NVIDIA does not publish the Tensor Core count. The Workstation Edition datasheet lists the generation only, while the number 752 appears in the StorageReview writeup and carries no official confirmation. RT cores number 188, which matches the Server Edition page. Ada publishes everything openly: 568 Tensor, 142 RT.

600 W: what capping the card actually costs

The doubled TDP scares procurement more than it should. The same card was run at a 600 W and a 300 W cap on llama.cpp:

Metric 600 W 300 W
Llama 3.1 8B Q8, prompt reading 14,040 tok/s 10,093 tok/s
Llama 3.1 8B Q8, generation 165.1 tok/s 155.6 tok/s
Llama 3.3 70B Q8, prompt reading 1734.6 tok/s 906.6 tok/s
Llama 3.3 70B Q8, generation 20.5 tok/s 19.3 tok/s
Peak TFLOPS at the cap 423.2 260.2

Source: 600 W versus 300 W measurements at Level1Techs

Halving the power cap costs roughly 6% on generation and close to half the throughput on reading a long prompt. In chat the difference barely registers. In RAG, agent pipelines and long-document work it shows up immediately, because there the prefill phase sets the response time. If the requirement is 96 GB inside a 300 W envelope, Max-Q covers it: the same GDDR7 with ECC in a 4.4 × 10.5 inch board, the RTX 6000 Ada form factor.

What Ada could not do at all

MIG carves the card into isolated instances: four at 24 GB, two at 48, or one across all 96. Ada has no hardware partitioning whatsoever, and the only route there is vGPU with a ceiling of 32 virtual machines per card (16 in mixed-size mode), while enabling vGPU shuts off the physical display outputs of the RTX 6000 Ada. That restriction is a footnote in its own datasheet.

MIG on Blackwell has an unpleasant history. Retail cards shipped with firmware 98.02.52.00.02, MIG needs at least 98.02.55.00.00, and the NVIDIA position was that the vBIOS update has to come from the partner that sold the card. Ask for the firmware revision before you sign off on the order. Separately, in Proxmox vGPU is officially supported on Server Edition only, and Workstation Edition drops out of that scenario.

What the benchmarks show

Test on RTX PRO 6000 Workstation Result
LM Studio, Llama 3.1 70B 31.84 tok/s
LM Studio, GPT-OSS 120B 163.1 tok/s
Stable Diffusion XL FP16 5.364 s per image
Stable Diffusion 1.5 INT8 0.395 s per image
Blender 4.4, Monster scene 7870 samples/min
V-Ray 12,128 vpaths
System draw under SDXL 918.5 W average, 1036.3 W peak

The power numbers were taken at the wall for the whole test bench, not the card alone. Source: StorageReview

Now the honest weak spot in every comparison of these two cards. Paired runs, where both boards went through the same build on the same day, are thin on the ground, and nearly all of them cover rendering rather than LLM work. StorageReview does have such pairs: Blender 4.4 on the Monster scene gives 7870 samples/min against 5633 for the RTX 6000 Ada, while V-Ray returns 12,128 vpaths against 10,766. That works out to 40% in one test and 13% in the other on the same bench, which is the most honest answer to “how much faster”: it depends on the workload more than on the generation. One more hard data point for Ada sits in The Register review: FLUX.1 Dev in BF16, 50 steps at 1024×1024 takes 37 seconds on it, and a fine-tune of Llama 3.2 3B over a million tokens about 30 seconds.

What the slide decks leave out

NVLink is never mentioned in the RTX PRO 6000 documentation. Third-party sources say it was removed and everything runs over PCIe Gen5, which again means two cards will not present a single 192 GB pool. As for ECC on GDDR7, users could not turn it off: on earlier generations it was disabled for render work, here the nvidia-smi commands have no effect. No official NVIDIA answer ever appeared in that thread, and nobody stated how many gigabytes ECC takes either.

Early drivers added work. Version 575.51.03 returned “No devices were found”, because Blackwell runs only on the open kernel module. Then came the Xid 119 GSP timeout thread, where both cases ended in an RMA, and separately Xid 79 with the card falling off the bus as late as 580.159.03. A pair of cards on certain motherboards hung the system within 5 to 60 minutes on PCIe Gen4, stability returned at Gen3, and the fix turned out to be enabling PCIe Spread Spectrum. The card is not a bad card. Just budget time for matching firmware, driver and motherboard, and do not schedule the install for the Friday before a delivery date.

When the upgrade pays for itself, and when Ada stays the sane choice

Buy Blackwell if the model runs from 70 billion parameters up, or 30B with a long context; if you need MIG to divide one card between teams; if you have a video pipeline and four ninth-generation NVENC engines cover the actual volume; if the workload reads long prompts constantly. And if the workstation already has a power supply with headroom, a chassis for the longer board and clearance behind it for airflow.

Stay on Ada if the model already lives inside 48 GB and is not growing; if the fleet is standardized and identical drivers are worth more to you than the extra throughput; if the power budget for the room is fixed. Going from 300 W to 600 pulls in a new supply, sometimes a new chassis, and in a multi-card build a dedicated 20 amp circuit as well. One more thing: neither generation has meaningful FP64, so for numerical simulation the move to Blackwell changes nothing.

What we have in stock

FAQ

What do two cards give you over one?

Not 192 GB. There is no NVLink, and the memory does not merge into one pool. The real benefit: keeping several specialized models resident at once, or parallelizing render jobs.

Does MIG work out of the box?

On retail cards at launch it did not. Firmware 98.02.55.00.00 is the minimum, and cards were shipping with 98.02.52.00.02. Working configurations were logged on the forum at 98.02.81.00.07.

Why does the card hold 450 W instead of 600?

The cap is set by the sense configuration of the 12VHPWR cable. In Supermicro servers one cable part number is configured for 450 W, and 600 W needs a different one. In nvidia-smi it reads as Max Power Limit 450 W with a minimum of 300.

Is a 15 amp circuit enough?

For one card, yes. For a multi-card build, no: a workstation with four Max-Q boards measured 1650 to 1800 W at the wall under full load, and the practitioner advice is a 20 amp circuit.

Can ECC be disabled for render work?

On earlier generations it could. On GDDR7 users report that the nvidia-smi commands do nothing, and no official explanation from NVIDIA has appeared in the relevant thread.

One RTX PRO 6000 or two DGX Spark units?

That is a choice between 96 GB of fast GDDR7 and 256 GB of slow unified memory. On one LMSYS bench with gpt-oss 20B the card produced 215 tokens per second of generation, Spark 49.7.

Not sure 96 GB is what you need?

Tell us which models or scenes you plan to run, how many people will share the card and what is already installed in the workstation. An engineer will work out whether RTX 6000 Ada covers it or Blackwell is the smarter buy up front, and will spec the power supply and chassis to match.

Get a consultation