Quantized weights17.03 GB+ 0.68 GB format overhead
strongcodingtool-uselong-context
Qwen2.5-Coder 32B Instruct
Mature coding model with disclosed architecture; individual benchmark evidence remains context-specific.
Model profile
Parameters32.5B32.5B active
Maximum context131,072
LicenseApache-2.0
Released2024-11-12
32K / one-agent memory trace
Using GGUF Q4_K_M. This is fit math, not a fixed “billions to GB” lookup.
KV cache8 GBarchitecture-aware
Total with workspace + safety30.23 GB56 GB recommended RAM
KV cache uses layers × KV heads × head dimension × K/V × cache precision × context × concurrency.
Architecture metadata
| Layers | 64 |
|---|---|
| Attention / KV heads | 40 / 8 |
| Head dimension | 128 |
Evidence records
inference
Measured dense Qwen 32B Q4 on M3 Ultra; model variant and memory capacity differ, so this is comparable evidence only.