Quantized weights36.99 GB+ 1.48 GB format overhead
unrankedgeneraltool-uselong-context
Llama 3.3 70B Instruct
Included as a general high-capability comparison, not claimed to be a coding-specialist winner.
Model profile
Parameters70.6B70.6B active
Maximum context131,072
LicenseLlama 3.3 Community License
Released2024-12-06
32K / one-agent memory trace
Using GGUF Q4_K_M. This is fit math, not a fixed “billions to GB” lookup.
KV cache10 GBarchitecture-aware
Total with workspace + safety57.32 GB96 GB recommended RAM
KV cache uses layers × KV heads × head dimension × K/V × cache precision × context × concurrency.
Architecture metadata
| Layers | 80 |
|---|---|
| Attention / KV heads | 64 / 8 |
| Head dimension | 128 |
Evidence records
No cataloged benchmark is sufficiently comparable. The model remains unranked on unsupported dimensions.