The Intelligence Factory Open the visualizer
Illustration of an AI campus at blue hour: a lit substation fed by a transmission line from the left, two long data halls with dry coolers on their roofs, a generator yard and farmland around it.

Power · Data · Heat · Six scales

The Intelligence Factory

Electricity in. Intelligence out.

Follow the power from the plants on a 2,000 km grid into one AI campus: 345,000 volts at the fence, down through the substation, the halls, a rack and a tray, to 0.8 V at a GPU die, and out as tokens. Switch to the data layer to follow the links back out, or to heat to follow every watt to the sky. The figures on each part and chart carry a label for what backs them: a published spec, a vendor claim, a third party's report, or this page's own math or assumption.

Personal educational project based on cited public sources. Not an official publication of my employer or the companies discussed. Models are schematic; estimates and assumptions are identified.

Illustration, drawn by an image model from this site's campus layout

01Visualizer

Fly through one campus at six scales, from a regional grid to a GPU die, and four more inside the links: a pluggable optical module, a co-packaged switch, a coherent module and copper cables. Every pin is a part with its own numbers in power, data and heat, each labeled by what backs it.

1.1 · Scenario

Your campus

100 MW · GB200 · 415 V AC · Warm waterChange it in the visualizer’s Scenario pane →

These are this model’s estimates for this scenario, not figures for a real building, even when a real campus is picked. Every chart on this page follows it.

02Power

From the meter to the die: where every megawatt goes and why the voltage steps down.

The regional map names real campuses by their own power source — Meta's Hyperion runs on new gas-fired plants (see it on the map) — and this campus has its own perimeter security and operations center watching it.

2.1 · The ledger

Where 100 MW goes

Start with 100 MW at the campus meter. Each row takes off what one stage turns into heat, spends on cooling, or hands to silicon other than the GPU. The highlighted rows belong to the scale you are looking at. The grid already lost about 5% before the meter.

Every row is this model’s estimate for your scenario. Click a row’s label to see its calculation and the published figures it starts from.

Conversion loss, becomes heat Cooling and building Network: NVLink, NICs, switches, optics Other silicon: CPUs, memory, drives

2.2 · Voltage

Volts down, amps up

Power is volts times amps. Each step down in voltage raises the current a conductor has to carry, which is why the last step happens millimeters from the chip.

03Data

Copper inside the rack, light across the hall, fiber between campuses: how fast each link is, what carries it, and how a model is split to fit.

Shared storage and the meet-me room where outside networks cross-connect each ride a network of their own, apart from the GPU fabric.

3.1 · Bandwidth

Bandwidth down, latency up

The data path has its own staircase. The closer two pieces of silicon sit, the faster they can talk. That is why a model is split the way it is: the chattiest work stays inside a rack on NVLink, less of it crosses the hall, and as little as possible crosses the country.

3.3 · Parallelism

How a model is split

The bandwidth staircase decides the layout. The chattiest kind of splitting goes where the links are fastest, and each step outward carries less, less often. Switch the 3D view to Data and open the hall to see one layout drawn on the rack tops.

04Heat

Every watt ends as heat. The path from a GPU die to the outside air, and what it costs in power and water.

Air-sampling fire detection watches the hall continuously, so it can find smoke before a fire grows.

4.1 · Temperature

Hot to cold

Every watt ends as heat, and heat only flows downhill. Each hop from die to outdoor air costs a few degrees, which is why warmer water is the trick: NVIDIA's newest spec feeds racks at 45 °C, so even a 35 °C afternoon can take the heat through dry coolers, without chillers.

05Tokens

What the whole machine makes, and what each token costs in energy, carbon and water, with the parts list it takes.

Open the Tokens card in the visualizer and click Show the math to see one decode step worked out as a single matrix-vector multiply.

5.1 · Cost

What a kilowatt-hour buys

Energy per token depends on more than the chip. Idle capacity, the building, the grid's carbon, the water on site and the training run behind the model all land on the bill. The campus comes from your scenario; set the rest here.

2,000

The number above is an illustrative serving assumption, scaled by memory bandwidth for the accelerator chosen above — not a measurement. The DeepSeek preset is a published decode (answer-writing) rate for GB200 specifically, so it’s off on other hardware.

60%

Share of GPU time spent producing tokens. Fleets keep idle capacity for peaks and failover; nobody publishes a fleet figure, so this is a scenario.

370 g

On-site water comes from the cooling choice above. It excludes water at the power plant, ≈1.8 L/kWh for typical thermal generation (NREL).

50 GWh
1,000 T tokens

A scenario, not a measurement: nobody publishes both numbers for one model. Training matters most for models that are trained big and used little.

–joules per token, all in
–tokens per kWh
–of that energy is training
–Wh per 500-token reply
–g CO₂ per reply
–mL water per reply, on site
  • Google, median Gemini text prompt (2025). Includes idle capacity, host CPU and DRAM, and PUE overhead; excludes training. Google's own write-up doesn't mention embodied hardware, so this page doesn't count that as included0.24 Wh · 0.03 g · 0.26 mL Reported
  • Llama 3.1 405B pretraining (Meta model card), location-based39.3M H100 GPU-hours · 11,390 t CO₂e Spec
  • GPT-4 training electricity (Epoch AI reconstruction)≈52–62 GWh Reported
  • Embodied carbon, HGX H100 board of 8 GPUs (NVIDIA)1,312 kg CO₂e Spec
  • LLaMA-65B on A100, GPU only (Samsi et al., 2023)3–4 J/token Reported
  • DeepSeek-R1 decode on GB200, per GPU (SemiAnalysis)≈13k tok/s Reported
  • NVIDIA: GB200 NVL72 vs H200, tokens per MW≈8–10× Vendor

5.2 · Inventory

What it takes: a 100 MW campus, counted

Sized from the same assumptions as the ledger: 100 MW at the meter, PUE 1.2, 132 kW racks. Real campuses differ in redundancy and layout; the counts are here to give a sense of scale.

How this site is built

How it was made

The page is written in TypeScript, drawn with Three.js (version 0.183.2) — which describes itself as a “JavaScript 3D Library” — and bundled with Vite (version 8.3.1), which calls itself “a build tool that aims to provide a faster and leaner development experience for modern web projects.” The build is static and deployed to GitHub Pages.

Most of the hardware you see was built in Blender — “Free and Open Source software, forever,” in its own words — by this project’s own Python build scripts rather than modeled by hand, and exported as glTF/GLB, the Khronos Group’s “royalty-free specification for the efficient transmission and loading of 3D scenes and models by engines and applications.” Twenty-five of those files, built from this project’s own Blender sources by its own scripts, cover the racks, trays, compute chips, campus buildings and vehicles, and the optical and co-packaged-optics modules; everything else in the scene is procedural, built by code from the model’s numbers, so it recounts and resizes itself when you change the scenario.

Every figure on the page is labeled by what backs it: Spec, Vendor, Reported, Calc. or Assumed, with its sources one click away. The Evidence page lists every figure that way, and the Method page lists every calculation and assumption behind them. Underneath, a scenario engine turns four choices — campus size, accelerator, power path and cooling — into every number, chart and sentence on the page.

Automated checks run before anything ships: Vitest unit tests (1,350 tests across 37 files, as of this writing) cover the engine, the content and the links; Playwright drives a real browser through checks named for what they look for — cycle (every scenario, scene and layer, watching for errors), views (every part framed clear of overlays), coplanar (flush surfaces that would flicker), parts (every numbered part has a pin) and ui (the page’s controls stay out of each other’s way, on phone and desktop alike). A quality governor adapts the 3D rendering to the device it’s running on, with Max quality, Auto quality and Battery saver settings for when you want to choose yourself.

The code, 3D models and automated checks were written with AI coding assistants, mainly Anthropic’s Claude and OpenAI’s Codex. I decided what to show, set the rule that every figure comes from a cited public source or a labeled assumption, and reviewed the work, sending back what was wrong or unsupported. The Evidence page and the automated checks let anyone verify the result.