GPU
The massively parallel compute engine
- Ada Lovelace, Blackwell, RDNA architectures: what actually changes from one generation to the next.
- CUDA cores for general-purpose compute, Tensor cores for the matrix multiplications behind neural nets.
- FP16, BF16 and FP8 precision: trading raw speed against numerical stability.
- PCIe 5.0 and NVLink: interconnect bandwidth dictates how well multi-GPU scales.
| RTX 4090 | ≈ 330 TFLOPS |
|---|---|
| RTX 5090 | ≈ 420 TFLOPS |
| H100 SXM | ≈ 990 TFLOPS |