On this page
The short answer: we have estimates, not a tested rig
Using 1.05 trillion total parameters from Mistral’s model documentation, raw 4-bit weights take about 525 GB, or 489 GiB. That is a storage calculation. It is not a promise that a machine with that much VRAM can serve the model.
Le Chonk uses a mixture-of-experts architecture. Only part of the model is active for a token, but the other experts still have weights. Depending on the runtime, those weights may live on GPUs, in system memory or in another storage tier.
Do not budget memory from active parameters alone. The release says 49B active; the model docs say 52B. Neither figure replaces the total weight count when estimating raw storage.
There is no locally tested hardware recommendation on this site yet. We have not verified downloadable weights, a quantized artifact or a supported inference framework for the preview.
How much room do the weights take?
The lower-bound formula is total parameters × bits per parameter ÷ 8. Here is the result for 1.05T parameters. The figures assume every parameter uses the selected width, without extra storage.
| Precision | Decimal GB | Binary GiB | Important limit |
|---|---|---|---|
| 16-bit · BF16 / FP16 | 2,100 | ≈ 1,956 | No runtime memory included |
| 8-bit | 1,050 | ≈ 978 | No quantization metadata included; format availability not verified |
| 4-bit | 525 | ≈ 489 | No quantization metadata included; format availability not verified |
GB and GiB are different units
One GB is 1,000,000,000 bytes. One GiB is 1,073,741,824 bytes. The same byte count looks smaller in GiB. Keep the units consistent when comparing this table with GPU specifications or operating-system reports.
Lower precision is not a free upgrade
Reducing bit width lowers the raw storage estimate. It does not prove that a supported quantization exists or that the resulting model meets your quality needs. A real artifact can use mixed precision and add scales, zero points, padding and other metadata.
Try a different memory estimate
Change the total parameter count or bit width. The calculation stays in your browser. It does not infer a GPU count or include architecture-specific KV cache.
Estimated weight storage
A theoretical lower bound. Add memory for KV cache, runtime buffers, quantization metadata and communication. This is not a tested GPU configuration.
What else needs memory?
KV cache
Attention state grows with context and concurrent requests. Its exact size depends on the architecture, cache precision and serving configuration.
Runtime buffers
The framework needs working space for activations and kernels. Peak use can differ from steady-state memory.
Quantization metadata
Low-bit formats may store scales, groups and higher-precision tensors. Raw bit arithmetic does not include them.
Parallel communication
Multiple devices need buffers for exchanging data. The parallel strategy can also replicate some tensors.
System RAM is a separate budget. It may hold files, loaded tensors and offloaded experts. Storage capacity, RAM and VRAM are not interchangeable, and a successful load does not prove usable response speed.
We do not publish an exact KV cache figure because the necessary architecture and serving details have not been verified here. A percentage added to the weights can be a planning allowance, but it is not a measured memory requirement.
Can offloading or more GPUs help?
CPU and expert offloading
A runtime may keep some weights in system RAM and move them when needed. This can reduce GPU residency. It can also make host-to-device bandwidth, CPU speed and scheduling the bottleneck. We have no measured throughput for such a Le Chonk setup.
Using several GPUs
Adding up the capacities is only a first check. The framework must support the model and a parallel layout that fits on every device. Interconnect speed, tensor replication, uneven allocation and the largest individual tensor can still block a setup.
For example, enough total capacity for raw weights does not leave space for runtime buffers or prove that the weights can be split in the needed way. Do not turn a division of GB by GPU capacity into a tested configuration.
Offloading to storage
Storage can hold large files, but moving weights from a drive is a different performance trade-off. Check whether the chosen runtime supports it and measure cold starts and steady-state speed before treating it as a usable serving plan.
Before buying or renting hardware
- Wait for a real weight download and read its license.
- Check the artifact’s file sizes and quantization format.
- Find an inference framework version that explicitly supports the architecture.
- Check GPU and host-memory requirements from a reproducible run.
- Start with one short request; then test your intended context and concurrency.
- Record peak memory, load time, throughput and output quality.
The hardware budget should follow the verified artifact and runtime. A trillion-parameter headline alone is not enough to choose a machine.
Follow the local setup readiness checklist →Sources & calculation notes
- Mistral Large 4 release announcement
API preview, planned weight release, 49B active parameters and reported benchmark results.
- Mistral Large model documentation
1.05T total, 52B active, 1M context. Displays both $0.68 / $2.09 and higher $1.36 / $4.18 prices. No discount end date confirmed.
Estimates use the model documentation’s 1.05T total. Formula and unit conversions are public in this page’s calculator. No GPU measurements were made by this site.