Tesla turns on 10k-node Nvidia H100 Cluster

Tesla’s plan to run a supercomputer built from 10,000 Nvidia H100 GPUs is prompting debate over how substantial this capability really is, both in scale and in cost, compared to leading cloud and HPC installations. Commenters question whether this undermines or complements Tesla’s in-house Dojo hardware effort, and whether Musk’s public claims around custom chips and Full Self-Driving timelines are realistic or largely promotional. Others focus on technical details such as FLOPS metrics, networking, and power requirements, while also criticizing the thin sourcing of the original news reports, which lean heavily on a small number of enthusiastic tweets.

Scope and Definition of “Dojo” vs H100 Cluster

  • Confusion over terminology: some use “Dojo” to mean all of Tesla’s ML training compute; others insist Dojo specifically refers to the custom-chip supercomputer.
  • Cited material indicates Tesla has both: a Dojo system and a separate 10,000‑H100 training cluster.
  • Some argue the “Dojo supercomputer” isn’t real until large-scale custom hardware is live; others say chips are in production at TSMC with thousands on order, so calling it vaporware is inaccurate.

Scale, Cost, and Infrastructure

  • 10,000 H100s are described as a very large cluster, in the same order of magnitude as top-listed supercomputers; large enough to pretrain 100B+ parameter LLMs.
  • Compared with historical crypto mining operations (e.g., 150k older AMD GPUs), 10k feels smaller in count but much bigger per-node and system-wise.
  • Commenters note that the $300M GPU figure excludes major costs: storage, RAM, networking (400–800 Gbit links), power, cooling, racks, data center buildout, and vendor engineering support.
  • Supply is characterized as Nvidia‑bottlenecked; suggestion that large buyers need relationships to secure hardware and NICs.

Reliability of Reports and Hype Level

  • Several see TechRadar/Tom’s Hardware coverage as second-hand “blogspam” built off a single tweet and Tesla-friendly social media.
  • Skepticism about marketing claims like “20x–30x performance,” and about whether 10k H100s are actually installed vs ordered or planned, given future-tense language and global H100 supply limits.
  • Others defend the social media sources as typically accurate Tesla/SpaceX news channels, but critics argue that tweets alone are weak evidence of deployment.

Technical and HPC Details

  • Discussion over perf metrics: article and Nvidia marketing lead with FP64 PFLOPS; unusual because consumer GPUs often have weaker FP64.
  • Nitpicks that INT8 ops are not “FLOPS”; should be counted as integer ops or “TOPS.”
  • Some debate why FP64 and FP32 throughput would be equal; speculation that this might imply design tradeoffs that “hamstring” FP32.

Broader Context: Tesla, Musk, and FSD

  • Debate over whether Tesla can out-build AI hardware compared to richer, more established AI players.
  • Mixed views on Musk: praised for enabling reusable rockets and popularizing EVs; criticized for overpromising (e.g., robo‑taxis/FSD timelines), management style, and public behavior.
  • Jokes that after all this compute, the conclusion might be that FSD still needs lidar.