strongcodingtool-uselong-context

Qwen2.5-Coder 32B Instruct

Mature coding model with disclosed architecture; individual benchmark evidence remains context-specific.

Model profile

Parameters32.5B32.5B active
Maximum context131,072
LicenseApache-2.0
Released2024-11-12

32K / one-agent memory trace

Using GGUF Q4_K_M. This is fit math, not a fixed “billions to GB” lookup.

Quantized weights17.03 GB+ 0.68 GB format overhead
KV cache8 GBarchitecture-aware
Total with workspace + safety30.23 GB56 GB recommended RAM
KV cache uses layers × KV heads × head dimension × K/V × cache precision × context × concurrency.

Architecture metadata

Layers64
Attention / KV heads40 / 8
Head dimension128

Evidence records

inference

Measured dense Qwen 32B Q4 on M3 Ultra; model variant and memory capacity differ, so this is comparable evidence only.

comparable-measurement
mlx-lm