FP8, FP4 and the shrinking bit: how number formats decide which GPU can run your model

Every weight in an AI model is a number, and the number of bits used to store it decides memory, speed, energy and which GPU you need. A plain-language guide to FP32, BF16, FP8, FP4 and block scaling, the hardware that supports each, and what it means for buying GPUs in Bangladesh.

CompTech Team
  • AI Infrastructure
  • GPU
  • LLM
  • Data Centre
Close-up of an NVIDIA H100 SXM module, the first data-centre GPU with FP8 tensor cores

A large language model is, underneath everything, a very large list of numbers. Llama 3.3 70B has 70 billion of them. DeepSeek-V3 has 671 billion. How many bits you spend on each of those numbers decides how much memory the model needs, how fast it answers, how much electricity it burns, and which GPU can run it at all.

Five years ago the answer was simple: 16 bits per number, and buy enough GPUs to hold them. Today the leading edge is 8 bits for most work and 4 bits for inference, and the newest GPUs from NVIDIA and AMD are built around those formats. This post explains what the formats are, why fewer bits usually work, how each generation of hardware fits in, and what that means if you are specifying GPUs for a project in Bangladesh.

What a floating-point number actually is

Every format is a trade between three things that share a fixed number of bits: a sign, an exponent and a mantissa. The exponent sets the range, how large or small a number can be. The mantissa sets the precision, how finely the values between two powers of two are spaced.

Bar diagram of the bit layout of FP32, TF32, BF16, FP16, FP8 E4M3, FP8 E5M2, FP6, FP4, INT8 and INT4, showing sign, exponent and mantissa widths with the maximum value and approximate precision of each
The common formats side by side. Each step down roughly halves the memory per value. Notice that BF16 keeps the full 8-bit exponent of FP32 and gives up precision instead, which is why it replaced FP16 for training.

Three lessons come straight from the diagram:

  • FP32 is a luxury. Seven decimal digits of precision is far more than a neural network needs. Training frameworks keep an FP32 "master copy" of the weights for the optimiser, but almost all the arithmetic happens in smaller formats.
  • BF16 won because of range, not precision. Google's Brain Float 16 keeps the FP32 exponent, so gradients never overflow, while FP16 tops out at 65,504 and needs loss scaling to survive. Every modern accelerator does BF16 at full speed.
  • FP8 comes in two flavours because 8 bits is not enough for one layout to do both jobs. E4M3 (4 exponent bits, 3 mantissa bits) is more precise and is used for weights and activations. E5M2 has the range of FP16 and is used for gradients during training. Both are defined in the Open Compute Project 8-bit floating point specification, based on the format NVIDIA, Arm and Intel proposed in 2022.

INT8 and INT4 are different animals. They have no exponent, so the spacing between values is fixed, and they rely entirely on an external scale factor. They were the first low-bit formats to get hardware support, and INT4 weight-only quantisation (GPTQ, AWQ, the GGUF files used by llama.cpp) is still the most common way to run a model on a laptop.

The trick that makes 4 bits usable

An FP4 value has four bits, which means it can take exactly sixteen values: 0, ±0.5, ±1, ±1.5, ±2, ±3, ±4 and ±6. On its own that grid is useless for weights that cluster around 0.01. The answer is to store a shared scale for a small block of neighbouring values, so that the sixteen-point grid is stretched or shrunk to fit each block's actual range. The GPU applies the scale inside the tensor core, so the matrix multiply still runs at full 4-bit speed.

Diagram comparing MXFP4, with 32 four-bit values sharing one 8-bit power-of-two scale, against NVFP4, with 16 four-bit values sharing one FP8 scale plus a per-tensor FP32 scale
Two ways to share a scale. The OCP Microscaling (MX) formats use 32-value blocks and a power-of-two scale. NVIDIA's NVFP4 uses 16-value blocks and a fractional FP8 scale, costing a quarter of a bit more per value and losing noticeably less accuracy.

The two block formats you will meet in practice are:

  • MXFP4, MXFP6 and MXFP8, from the OCP Microscaling specification backed by AMD, Arm, Intel, Meta, Microsoft, NVIDIA and Qualcomm. A block of 32 values shares one 8-bit exponent-only scale. OpenAI's open-weight gpt-oss models ship their expert layers in MXFP4, and AMD's CDNA 4 GPUs run MX formats natively.
  • NVFP4, introduced with NVIDIA Blackwell in 2025. A block of 16 values shares one FP8 E4M3 scale, and a single FP32 scale covers the whole tensor. The smaller block and fractional scale track the data more closely. NVIDIA's published results for DeepSeek-R1 quantised to NVFP4 show an MMLU score of 90.7 percent against 90.8 percent for the FP8 original.

The same idea works at 8 bits. DeepSeek-V3, the first frontier-class model trained mostly in FP8, scales activations in tiles of 128 values and weights in 128 by 128 blocks, promoting partial sums to FP32 every few steps. Its technical report puts the full training run at 2.79 million H800 GPU-hours, a fraction of what BF16 training of a comparable model would have cost.

Where each format is used

Training. The standard recipe since 2022 is BF16 for everything with FP32 master weights, and the standard upgrade on Hopper or newer hardware is to run the big matrix multiplies in FP8 through NVIDIA's Transformer Engine or AMD's equivalent in ROCm. That roughly doubles throughput at the cost of some engineering care. NVIDIA Research has shown a 12-billion-parameter model pretrained on 10 trillion tokens in NVFP4 matching its FP8 baseline, so 4-bit training is coming, but in 2026 it is still research rather than something you would bet a production run on.

Fine-tuning. QLoRA, the method behind most low-cost fine-tuning, freezes the base model in a 4-bit format (NormalFloat 4) and trains small adapter matrices in BF16. That is what lets a 70B model be adapted on a single 80 GB GPU.

Inference. This is where low precision pays off most, and in three places:

  1. Weights. FP8 is now the default for serving large models: Meta publishes Llama 3.1 405B in FP8, and DeepSeek-V3 and R1 are FP8 natively. FP4 halves the memory again on Blackwell and CDNA 4.
  2. Activations. Quantising the values flowing between layers as well as the weights is what lets the tensor cores run at their FP8 or FP4 rate. Weight-only quantisation saves memory but computes in BF16.
  3. KV cache. The memory a model uses to remember the conversation so far grows with every token. Storing it in FP8 doubles the number of concurrent users a GPU can serve, often a bigger win than the weights.

What it buys you

Bar chart of GPU memory needed for the weights of six models from 8 billion to 671 billion parameters in BF16, FP8 and FP4, with dashed lines for the memory of one H100, H200, B200 and MI355X
Memory for the weights alone. A 70B model drops from two H100s to one in FP8, and fits on a 48 GB card in FP4. DeepSeek-R1 needs a full eight-GPU H200 server in its native FP8; in NVFP4 the weights fit in two B200s, and four serve it comfortably.

Memory is the obvious gain, and it decides how many GPUs you must buy before you can serve a model at all. A 70B model in BF16 needs 140 GB and therefore two H100s linked by NVLink. In FP8 it needs 70 GB and one H100, with little room for context. On an H200 or a 96 GB RTX PRO 6000 it fits comfortably. As 4-bit weights it fits on a 48 GB L40S with headroom, though that card converts them to FP8 to compute.

Speed follows memory more closely than most people expect. When a model generates text one token at a time, each token requires reading every weight from memory once. At small batch sizes the GPU is waiting on memory bandwidth, not arithmetic, so halving the bytes roughly doubles tokens per second. A 70B model in FP8 on an H200 with 4.8 TB/s of bandwidth cannot exceed about 68 tokens per second per sequence; in FP4 the ceiling rises to about 120. At high batch sizes the arithmetic rate matters too, and that is where the FP8 and FP4 tensor cores earn their keep.

Energy scales with both. Fewer bits moved and fewer bits multiplied means fewer joules per token. NVIDIA's figures for GB200 NVL72 against an H100 cluster claim up to a 25-fold reduction in energy per token for large models at the same latency, most of it from FP4 and the larger NVLink domain. Independent measurements are less dramatic but point the same way.

Accuracy is the price, and it is now small if the format is chosen with care:

  • FP8 with per-tensor or per-block scaling is close to lossless for models above about 7B parameters. Most benchmarks move by a tenth of a point.
  • NVFP4 and MXFP4 with calibration typically lose under one percent on knowledge benchmarks for large models. Reasoning models and small models are more sensitive, and maths and code tasks show the gap first.
  • INT4 weight-only (GPTQ, AWQ) loses one to three percent at 70B and more at 8B. It is still the right choice for CPUs, Apple silicon and older GPUs where no FP4 hardware exists.
  • Always evaluate on your own tasks. A Bangla question-answering system may behave differently from an English benchmark.

Which GPU supports what

"Supports" has a precise meaning here: the tensor cores can multiply in that format at full rate. Any GPU can load FP8 or FP4 weights and convert them to BF16 on the fly, and software such as the Marlin kernels does this well on A100-class hardware, but it only gets the memory and bandwidth benefit, not the compute benefit.

NVIDIA A100 PCIe accelerator in its gold and black shroud
NVIDIA A100 (Ampere, 2020). It introduced TF32 and BF16 tensor cores and INT8 at full rate, but has no FP8. Many clouds in South Asia still run on it. Photo: NVIDIA, CC BY-SA 4.0.
Four NVIDIA H100 PCIe cards standing in a row on a marble surface
NVIDIA H100 (Hopper, 2022) brought the first FP8 tensor cores and the Transformer Engine that picks E4M3 or E5M2 layer by layer. The H200 is the same chip with 141 GB of faster memory. Photo: 极客湾 Geekerwan, CC BY 3.0.
An NVIDIA HGX B200 eight-GPU baseboard with its heat sinks, labelled HGX B200 NVL8
An HGX B200 baseboard with eight Blackwell GPUs. Blackwell added FP4 and FP6 tensor cores and the second-generation Transformer Engine with block scaling in hardware. Photo: Pokiiri, CC BY-SA 4.0.
A GB200 compute tray with copper cold plates over two Blackwell GPUs and a Grace CPU, shown at Computex 2024
A GB200 compute tray with its liquid cold plates. Seventy-two of these GPUs in one NVL72 rack give about 650 petaflops of dense FP4, which is the number behind most "AI factory" headlines. Photo: 极客湾 Geekerwan, CC BY 3.0.
Accelerator Year Memory Bandwidth Lowest float format Dense FP8 Dense FP4
NVIDIA A100 80 GB 2020 80 GB HBM2e 2.0 TB/s BF16 (INT8 624 TOPS) — —
NVIDIA L40S 2023 48 GB GDDR6 0.86 TB/s FP8 733 TFLOPS —
NVIDIA H100 SXM 2022 80 GB HBM3 3.35 TB/s FP8 1,979 TFLOPS —
NVIDIA H200 2024 141 GB HBM3e 4.8 TB/s FP8 1,979 TFLOPS —
AMD Instinct MI300X 2023 192 GB HBM3 5.3 TB/s FP8 2,615 TFLOPS —
AMD Instinct MI325X 2024 256 GB HBM3e 6.0 TB/s FP8 2,615 TFLOPS —
Intel Gaudi 3 2024 128 GB HBM2e 3.7 TB/s FP8 1,835 TFLOPS —
NVIDIA B200 2025 192 GB HBM3e 8.0 TB/s FP4 (NVFP4, MXFP4), FP6 4,500 TFLOPS 9,000 TFLOPS
NVIDIA B300 (Blackwell Ultra) 2025 288 GB HBM3e 8.0 TB/s FP4 ≈ B200 15,000 TFLOPS
AMD Instinct MI355X 2025 288 GB HBM3e 8.0 TB/s FP4 (MXFP4), FP6 5,000 TFLOPS 10,000 TFLOPS
NVIDIA RTX PRO 6000 Blackwell 2025 96 GB GDDR7 1.6 TB/s FP4 yes yes
NVIDIA GeForce RTX 5090 2025 32 GB GDDR7 1.8 TB/s FP4 yes yes

Vendor peak figures without sparsity, rounded. Marketing sheets usually quote the "with sparsity" number, which is exactly double. Google's TPUs and AWS Trainium follow the same path: Trillium and Trainium 2 top out at INT8 and FP8, Ironwood and Trainium 3 add FP8 at scale and, in Trainium 3's case, FP4.

NVIDIA GeForce RTX 5090 Founders Edition graphics card standing upright
The consumer GeForce RTX 5090 carries the same Blackwell FP4 tensor cores as the data-centre parts. With 32 GB it runs a 70B model in NVFP4 on a single desktop card. Photo: ZMASLO, CC BY 3.0.
Jensen Huang on stage at CES 2025 in front of a slide showing the RTX Blackwell GPU with 4,000 AI TOPS
NVIDIA's CES 2025 keynote. The "4,000 AI TOPS" headline for RTX Blackwell is an FP4-with-sparsity figure, which is why it is three times the Ada generation's FP8 number. Photo: Pronoia, CC0.

The software that does the work

Formats are useless without kernels, and the tooling has matured quickly:

  • NVIDIA: TensorRT Model Optimizer quantises a Hugging Face checkpoint to FP8 or NVFP4 with calibration; TensorRT-LLM, vLLM and SGLang serve the result. Transformer Engine handles FP8 training in PyTorch and JAX.
  • AMD: Quark quantises to FP8 and MXFP4; vLLM and SGLang run on ROCm with FP8 kernels for MI300 and MX kernels for MI350.
  • Intel: Neural Compressor and the Gaudi software stack cover FP8 on Gaudi 2 and 3.
  • Everyone else: llama.cpp's GGUF formats (Q4_K_M, Q8_0 and friends) are integer block formats that run on CPUs, Apple silicon and any GPU. They are the practical choice below the data-centre tier.
  • Model hubs: it is now normal for a model to be published in BF16, FP8 and FP4 variants on the same day, with the quantised versions coming from the model's own authors or from NVIDIA, Red Hat and Neural Magic.

What this means for Bangladesh

Most GPU capacity in the country today is A100 and older, bought for computer vision and traditional machine learning. For large language models that hardware is memory-bound and has no FP8, so each A100 serves roughly half the users of an H100 and a quarter of a B200 on the same model. Three practical conclusions follow.

For an inference service in 2026, FP8 is the floor. H100, H200, MI300X or Gaudi 3 class hardware runs the current open models in their native FP8 and uses FP8 KV cache to multiply concurrency. Buying anything older for a chatbot, a document-search assistant or a Bangla speech pipeline is paying for memory you will spend on padding.

For a sovereign AI cloud, FP4 changes the power equation. Bangladesh's binding constraint is reliable megawatts, not capital. A Blackwell or CDNA 4 rack serving in NVFP4 or MXFP4 delivers two to three times the tokens per megawatt of a Hopper rack in FP8. If a facility has 2 MW of IT load, that difference is the gap between serving one ministry and serving ten. It also halves the memory needed per model, which means more distinct models per rack: a Bangla LLM, a legal model, a medical model and a coding model side by side.

For local-language fine-tuning, BF16 plus FP8 matmuls is enough. Adapting a 70B base model to Bangla text on H100-class hardware with QLoRA or FP8 full fine-tuning is a job of days, not months, and it does not need the newest generation. Spend the money on data and evaluation instead.

CompTech has deployed and supported GPU infrastructure in Bangladesh for over a decade, from single-server inference boxes to multi-rack clusters. If you are sizing a GPU purchase and want to know which format and which generation your workload actually needs, talk to us. The answer is often fewer, newer GPUs than the first quote suggests.


Cover photo of an NVIDIA H100 by 极客湾 Geekerwan, CC BY 3.0. Other photographs are credited in their captions and used under their Creative Commons licences from Wikimedia Commons. Diagrams and chart by CompTech. Performance figures are vendor-published dense peaks as of October 2026; always check the current datasheet.