Skip to content

Resources · AI Infrastructure Lifecycle

Inside a retired AI rack

Written for the person who owns the racks rather than the person who buys disposition services. Everything below comes from the manufacturers' own published configurations — because the component census is the part of this that nobody disputes, and it is also the part that decides how the project should be run.

The premise

One rack, more than two hundred assets

Decommissioning projects are usually scoped from a chassis count, because a chassis count is what the asset register holds. For conventional infrastructure that approximation is survivable. For accelerated infrastructure it is not, and the reason is arithmetic rather than opinion.

A single current-generation rack contains ninety storage devices. The register will show eighteen compute trays.

What follows is the census, taken from published specifications, and what each line of it implies for inventory, sanitization and recovery.

Published specifications

Two reference systems, as the manufacturer documents them

System Form factor Weight Power Storage
NVIDIA DGX H100 / H200 8U rackmount 287.6 lb (130.45 kg) max 10.2 kW max 2 × 1.92 TB NVMe M.2 (RAID 1, OS) + 8 × 3.84 TB NVMe U.2 self-encrypting (RAID 0, data cache)
NVIDIA GB200 NVL72 Full rack — 18 compute trays, 9 NVLink switch trays, 8 power shelves Published by system builders in the 3,000 lb class, assembled ≈120 kW Per compute tray: 4 × 3.84 TB E1.S NVMe (RAID 0) + 1 × 1.92 TB M.2 NVMe boot

Source: manufacturer product documentation. Assembled rack weight varies by system builder and is published by the builder rather than the chip vendor, so it is stated here as a class rather than a single figure.

The census

One GB200 NVL72 rack, counted properly

72
Blackwell GPUs
18 compute trays × 4
36
Grace CPUs
18 compute trays × 2
18
NVSwitches
9 switch trays × 2, at 72 NVLink ports each
90
NVMe storage devices
18 trays × (4 × E1.S data + 1 × M.2 boot)
48
Power supply units
8 shelves × 6 × 5.5 kW, N+N redundant

Before counting memory modules, transceivers, cables, rails or the management controllers.

What it implies

Six consequences for how the retirement should be run

The drive count is the headline, not the GPU count

Ninety NVMe devices in a single rack, each of which held datasets, checkpoints or model weights. That is the actual data exposure, and it is also the only component class with a published, standards-backed sanitization path. An inventory taken at chassis level records eighteen line items where there are more than two hundred assets.

Some of those drives are self-encrypting

The data-cache drives in a DGX H100 are SED units. That matters operationally: where a drive supports it, cryptographic erase is a purge technique recognised by the sanitization standards, and it is dramatically faster than an overwrite pass on a 3.84 TB device. Knowing which drives in an estate are SED and which are not is a scoping question worth asking before a schedule is agreed.

The interconnect is a separate asset class

Eighteen NVSwitches, the InfiniBand adapters in every compute node, and the transceivers and cabling that connect them. Fabric hardware is fungible in a way accelerators are not — no matched-set constraint, no baseboard coupling — and a conventional test bench can actually power it.

Forty-eight power supplies is a real line item

Power shelves and PSUs are rarely inventoried in a decommissioning project and are routinely worth more than the effort of recording them. The same is true of rails, busbars and cable management, which are frequently written off as fixtures.

Weight and power are planning inputs, not footnotes

A DGX H100 chassis weighs 287.6 lb by the manufacturer's own specification, which is above what one person can safely handle and above the rating of a good deal of general materials-handling equipment. At rack scale the number climbs into the thousands of pounds, at which point the floor loading of the room and the load path out of it become part of the plan.

Some racks cannot be air-tested at all

The current generation of the densest systems is entirely liquid-cooled, with no air-cooled variant. That constrains what any facility can do with a tray after it arrives, and it is a large part of why grading in this category should be built on evidence captured while the equipment is still running rather than on a bench test afterwards.

FAQ

What project owners ask

Why does inventory granularity matter so much here?
Because the components inside one chassis have different sanitization requirements and residual values that differ by orders of magnitude. Recording a compute tray as one line simultaneously under-sanitizes — because the drives inside it are not individually accounted for — and under-recovers, because the memory, fabric adapters and processors are never separately valued. Neither error is visible at the time. Both surface later, one in an audit and one in a settlement statement.
How many data-bearing devices should we expect?
Far more than the chassis count suggests. Using the manufacturers' own published configurations: a DGX H100 carries ten NVMe devices — two OS drives and eight data-cache drives. A GB200 NVL72 rack carries ninety, five per compute tray across eighteen trays. Add the management controllers, which hold credentials, network configuration and logs rather than workload data but are still worth treating deliberately.
Is the accelerator memory a data-bearing component?
It is volatile memory, and volatile memory does not retain contents across a power-down, transport and warehousing cycle. The complication is evidentiary rather than physical: no published standard defines a sanitization procedure for it, so there is no verification step and no certificate anyone can honestly issue. We set that out in full in our analysis of what the standards actually cover.
What gets missed most often?
In our reading of how these projects are scoped: the boot drives, because they are physically small and sit behind the data drives in everyone's mental model; the management controllers, because they are not classified as data-bearing; the fabric transceivers, because they are counted as cabling; and the power shelves, because they look like infrastructure rather than assets. All four are recoverable, and three of the four carry configuration state.
Does any of this change how a project should be sequenced?
Yes, in one specific way that is worth planning for. A great deal of the information that determines what equipment is worth — error history, operating state, configuration — is readable while the system is still running and becomes difficult to obtain once it has been powered down and shipped. Sequencing that capture into the on-site work, rather than treating it as something the receiving facility will do later, is the single most useful scheduling decision available in this category.

Scoping a retirement against a real component census?

We inventory at component level because the register almost never does. Tell us the site and the deadline.