Quantized weights3.99 GB+ 0.16 GB format overhead
capablecodingtool-uselong-context
Qwen2.5-Coder 7B Instruct
A responsive entry-level coding model; capability is not directly comparable with larger agentic models.
Model profile
Parameters7.62B7.62B active
Maximum context131,072
LicenseApache-2.0
Released2024-11-12
32K / one-agent memory trace
Using GGUF Q4_K_M. This is fit math, not a fixed “billions to GB” lookup.
KV cache1.75 GBarchitecture-aware
Total with workspace + safety8.15 GB24 GB recommended RAM
KV cache uses layers × KV heads × head dimension × K/V × cache precision × context × concurrency.
Architecture metadata
| Layers | 28 |
|---|---|
| Attention / KV heads | 28 / 4 |
| Head dimension | 128 |
Evidence records
No cataloged benchmark is sufficiently comparable. The model remains unranked on unsupported dimensions.