Storing bfloat16 and FP8 tensors in HDF5 without a conversion tax
by Scot Breitenfeld, Director of Engineering, The HDF Group Training and deploying large models requires transferring significant volumes of low-precision tensors, including weights, gradients, activations, and KV caches. bfloat16 maintains float32’s exponent range, ensuring numerical stability with reduced storage requirements. FP8 E4M3 and E5M2 are standard for H100-class hardware, while FP6 and FP4 enable the […]