Can my hardware run it?
Mistral Large 4 has 1.05 trillion parameters, of which 52 billion are used for each token, according to its model card. Even compressed, the weights alone need hundreds of gigabytes of memory.
Source: Mistral's model cardMistral Large 4 has 1.05 trillion parameters, of which 52 billion are used for each token, according to its model card. Even compressed, the weights alone need hundreds of gigabytes of memory.
Source: Mistral's model card| Precision | Bits per weight | All weights | Read per token |
|---|---|---|---|
| BF16 | 16 | 2,100 GB | 104 GB |
| FP8 | 8 | 1,050 GB | 52 GB |
| 4-bit | 4 | 525 GB | 26 GB |
"Read per token" is the size of the active weights, which every generated token reads once. All sizes are lower bounds: parameters × bits per weight ÷ 8, in decimal gigabytes. Quantized files are a little larger because they also store scaling factors. Running the model needs more memory on top, for the context (the KV cache) and for the inference software. Mistral has not published the architecture details needed to estimate the KV cache yet.
Generating one token means reading every active weight from memory once. Dividing your memory bandwidth by the size of the active weights gives the most tokens per second a single request can reach. Compute, the KV cache and the links between GPUs all push the real number lower, so treat it as a ceiling, not a prediction.