CNEX / INFRASTRUCTURE FIELD GUIDE

NVIDIA GB300 NVL72 Power Plant Deployment Guide

From rack-scale architecture to facility readiness. A technical blueprint for deploying Blackwell Ultra inside an operating power plant.

By CNEX / Enterprise Infrastructure

ENGINEERED FOR THE PLANT

One rack. A new class of AI infrastructure.

Scope the complete system—not just the GPUs. Align compute, power, cooling, and operations before deployment.

72-GPU architecture and NVLink fabric

Power delivery and liquid-cooling readiness

OT integration and workload migration

01 / SYSTEM ARCHITECTURE

What is GB300 NVL72?

One rack. 72 GPUs. Acting as a single massive accelerator. Blackwell Ultra GPUs and Grace CPUs are connected by fifth-generation NVLink in a fully liquid-cooled, rack-scale system.

72 GPUs

Compute tray

18 compute trays integrate 72 Blackwell Ultra GPUs. Each tray combines four GPUs and two Grace CPUs.

Compute trays

18

GPUs per tray

4

Cooling

Cold plates

36 CPUs

Grace CPU

The CPU layer orchestrates data movement and connects the GPU complex to system memory.

CPU architecture

Arm

Cores total

2,592

CPUs per tray

2

~120 kW

Rack power

Power shelves feed the rack through a DC busbar. Plan the complete electrical and thermal envelope.

Power delivery

48 V DC

Facility supply

3-phase

Design basis

OEM-specific

02 / SPECS DEEP DIVE

Compute, memory, and fabric.

The architecture is designed for large-model inference and reasoning. Figures below follow the reference’s planning envelope; precision, configuration, and OEM implementation matter.

NVFP4

Compute

Blackwell Ultra brings native low-precision acceleration to demanding AI workloads.

Architecture

Blackwell Ultra

GPU count

72

FP4 per GPU*

~15 PFLOPS

Execution

Tensor Cores

HBM3e

Memory

High-bandwidth memory keeps larger model weights and active contexts close to compute.

HBM per GPU

288 GB

HBM total

~20 TB

GPU bandwidth

~8 TB/s

Memory stack

12-high

NVLINK 5

Fabric

A single NVLink domain reduces cross-device communication bottlenecks.

NVLink per GPU

1.8 TB/s

Rack bandwidth

~130 TB/s

NVLink domain

72 GPUs

Scale-out

ConnectX-8

03 / WHY IT MATTERS

Four reasons to bring AI inside the plant.

Power generation creates complex, sensitive, real-time data. Local compute places model execution closer to the systems and people that need it.

01 / MODEL CAPACITY

Plant-scale language models

Bring operating procedures, maintenance histories, P&IDs, and OEM documentation into a private knowledge environment. Large memory capacity supports demanding model deployments.

02 / SIMULATION

Real-time digital twins

Connect simulation and visual models with live plant data. High-bandwidth GPU communication supports complex, multi-model digital-twin pipelines.

03 / OPERATIONAL CONTEXT

Reasoning for operations

Investigate trips, alarms, and maintenance patterns with reasoning workflows. Keep operator review and plant safety controls separate from AI recommendations.

04 / DATA CONTROL

Sovereignty by design

Keep sensitive operational data within the site security perimeter. Apply access control, audit logging, and an approved OT/IT integration boundary.

04 / POWER PLANT TOPOLOGY

Four zones. One controlled data path.

Place the rack in a dedicated AI infrastructure room. Separate plant control, data aggregation, compute, and operator-facing applications with explicit trust boundaries.

OT DATA SOURCES

Zone 1 — Plant floor

DCS controllers • PLC cabinets • Vibration and thermal sensors • Edge modules

Interface / OPC UA / Industrial Ethernet

Boundary / Approved read-only extraction

CONTROLLED TRANSFER

Zone 2 — Data aggregation

Plant historian • Time-series buffer • DMZ firewall • OT/IT bridge

Transport / Private, encrypted link

Governance / Audit-logged data access

RACK-SCALE COMPUTE

Zone 3 — AI infrastructure

GB300 NVL72 rack • CDU • Switchgear • Scale-out fabric • Model registry

Facility envelope / ~120 kW rack planning basis

Cooling / Dedicated liquid loop

HUMAN-IN-THE-LOOP

Zone 4 — Applications & operations

Plant LLM console • Digital twin viewer • Reasoning workspace • Operations dashboards

Serving / Triton inference workflows

Control / Operator approval retained

05 / POWER & COOLING

120 kW per rack changes the facility equation.

This is not a standard server-room installation. Coordinate electrical capacity, heat rejection, structural loading, and service access as one engineering package.

ELECTRICAL

Power delivery

Validate upstream capacity, protection, redundancy, and the complete OEM power-shelf configuration.

Rack draw* / ~120 kW

Supply basis* / 480 V, 3-phase

Rack busbar* / 48 V DC

THERMAL

Liquid cooling

Direct-to-chip cooling requires a commissioned secondary loop and a facility heat-rejection strategy.

Cooling method / Direct liquid

CDU planning* / 150–200 kW

Acceptance / Flow / pressure / leak test

STRUCTURAL

Physical footprint

Survey the delivery route and structural design before the rack arrives. Preserve maintenance clearances.

Cabinet mass* / ~1.36 tonnes

Floor loading* / ~440 psf

Access / Front and rear service

CONNECTIVITY

Network fabric

Engineer scale-out connectivity alongside the OT boundary, management network, and storage paths.

NIC class* / 800 Gb/s

InfiniBand / Quantum-X800

Ethernet / Spectrum-X

06 / WORKLOADS

What runs on the rack.

Map each workload to a measurable operational objective. Consolidate inference where appropriate, then validate isolation, latency, and capacity under representative load.

01

Plant foundation LLM

Search procedures, maintenance records, incident reports, and OEM manuals in natural language. Ground answers in approved site documentation.

PRIVATE KNOWLEDGE

02

Real-time digital twin

Combine simulation and plant telemetry to explore equipment behaviour, visualise conditions, and evaluate operational scenarios.

SIMULATION

03

Reasoning agent

Investigate operational events using historical context and structured evidence. Present recommendations for engineer review.

DECISION SUPPORT

04

Multi-modal plant vision

Process thermal imagery, inspection footage, and acoustic signals to support detection of equipment anomalies.

VISION + LANGUAGE

05

Heat-rate optimisation

Serve combustion, vibration, and emissions models alongside one another. Test recommendations against site operating constraints.

MULTI-MODEL INFERENCE

06

Synthetic data & training

Use available capacity for fine-tuning and simulation-generated datasets. Track model lineage and promote through controlled releases.

MLOPS

07 / HARDWARE COMPARISON

From Hopper to Blackwell Ultra.

Compare system classes—not just GPU generations. HGX figures describe an 8-GPU node; NVL72 figures describe a complete 72-GPU rack.

Specification

H200 / HGX 8-GPU

B200 / HGX 8-GPU

GB200 NVL72

GB300 NVL72

Architecture

Hopper

Blackwell

Blackwell

Blackwell Ultra

GPUs per system

8

8

72

72

HBM per GPU

141 GB HBM3e

192 GB HBM3e

192 GB HBM3e

288 GB HBM3e

Total HBM*

~1.13 TB

~1.5 TB

~13.4 TB

~20 TB

NVLink per GPU

900 GB/s

1.8 TB/s

1.8 TB/s

1.8 TB/s

NVLink domain

8 GPUs

8 GPUs

72 GPUs

72 GPUs

Power envelope*

~10 kW / node

~14 kW / node

~120 kW / rack

~120 kW / rack

Cooling*

Air / rear-door HX

Air / liquid

Direct liquid

Direct liquid

Primary fit

Existing AI workloads

Enterprise AI nodes

Rack-scale AI

Large-scale reasoning

* Indicative reference figures, not a final engineering specification. Confirm performance precision, electrical supply, CDU sizing, floor loading, and cooling conditions with the selected OEM and qualified site engineers before procurement.

08 / DEPLOYMENT PATH

A phased path from survey to production.

An illustrative 16-week programme—not a delivery commitment. Procurement, permitting, utility readiness, and site acceptance determine the actual schedule.

WEEKS 1–3

Survey & design

Review floor loading, electrical capacity, coolant conditions, network routes, and OT integration. Agree the site acceptance criteria.

WEEKS 4–6

Facility preparation

Prepare switchgear, CDU placement, manifold runs, containment, and the approved network boundary.

WEEKS 7–12

Rack & platform integration

Position and commission the rack. Validate cooling, fabric, storage, model registry, and inference services.

WEEKS 13–16

Workload migration & handover

Benchmark priority workloads, rehearse recovery, train operators, and complete production acceptance.

09 / ENGINEERING FAQ

Before you commit to rack-scale AI.

Common questions for infrastructure, operations, and procurement teams.

SIZING

Do we need NVL72, or is an 8-GPU system enough?

Start with model size, concurrency, latency, and growth requirements. A smaller system may suit bounded workloads; NVL72 addresses larger scale-up domains and demanding multi-model deployments.

FACILITY

Can an existing IT room handle the rack?

Do not assume it can. Confirm available power, cooling-water conditions, floor loading, delivery access, and service clearances through a site engineering assessment.

COOLING

Is liquid cooling mandatory?

GB300 NVL72 is a liquid-cooled rack-scale platform. Plan the CDU, secondary loop, water quality, leak detection, and facility-side heat rejection together.

MIGRATION

What does migration from H200 involve?

Validate the software stack and model performance on the target platform. The larger change is often at the facility layer: power, cooling, networking, commissioning, and operational readiness.

START WITH SITE READINESS

Plan your plant’s next generation of compute.

Bring your plant layout, electrical capacity, and cooling-water specifications. Define what fits today, what needs to change, and a practical deployment path with CNEX.

72

72

Blackwell Ultra GPUs

~20 TB

~20 TB

Total HBM3e

130 TB/s

130 TB/s

NVLink fabric

~120 kW

~120 kW

Rack planning envelope

Building the future of AI infrastructure with unmatched speed and efficiency.

Keep in touch

Follow us

Powered by

CambridgeNexus

Building the future of AI infrastructure with unmatched speed and efficiency.

Keep in touch

Follow us

Powered by

CambridgeNexus

Building the future of AI infrastructure with unmatched speed and efficiency.

Keep in touch

Follow us

Powered by

CambridgeNexus