🖥️ Asel's AI Hardware Store NVIDIA GEAR

What is a "cluster"?

In one line: a cluster is several computers wired together closely enough that they act like one bigger computer — combining their memory, their power draw, and their price into one number.

The honest difference: real clustering vs. a pile of cards

Both options below can technically hold a huge model in memory. They are not the same thing.

NVIDIA RTX PRO 6000 Blackwell Server Edition

🗄️ A pile of desktop-class cards

Example: NVIDIA RTX PRO 6000 Blackwell Server Edition, bought in bulk and networked with standard PCIe/Ethernet.

Per card$13,000 / 96 GB / 600 W

These cards were never designed to work as a team. Each one is its own island; data has to hop between them over a relatively slow connection. It's the cheapest way to reach a big memory number, but it's slower and less reliable under heavy, sustained load.

Scaling up further: the full rack

NVIDIA GB300 NVL72

NVIDIA GB300 NVL72 takes real clustering to its extreme: 72 GPUs, all wired together with an NVLink switch fabric, pooling 20,000 GB of memory into what behaves like one enormous machine.

Power reality check: that rack draws 135.0 kW sustained — about 112 average homes' worth of continuous electricity, 24 hours a day.

When combining machines into a cluster, what changes?

MemoryAdds up — a cluster's total memory is the sum of every machine's memory, so bigger models fit.
PowerAdds up too — every machine you add, you pay for in watts as well as dollars.
PriceAdds up (plus, for real integrated systems, a premium for the engineering that lets them act as one).
SpeedDoes not simply add up — how fast a cluster works as a team depends entirely on how well the machines are wired together. This is the one number a pile of desktop cards can't buy its way around.