Here is the English transcription of the video @ https://x.com/RJDAIGOGO/status/2068160949606133955, which explains the technology behind High Bandwidth Memory (HBM).
00:00 This is the mainboard of a consumer gaming graphics card, the 5090. The most eye-catching part is the GPU chip in the middle.
00:06 Surrounding the GPU chip, we can see a ring of VRAM modules. The instructions and data needed for GPU operations, as well as temporary results, are stored here.
00:14 Now, this is a node from an NVIDIA GB300 AI data center. This node contains four GPUs, but if you look at their surroundings, you won't see any VRAM modules like the ones on the 5090. Does this mean AI server GPUs don't need VRAM to work?