Magistral Small 1.2 24B

24B

Mistral Magistral

Mistral's reasoning model on a Mistral Small 3.2 base, with reasoning wrapped in [THINK] tags. Q4 ~14GB puts explicit chain-of-thought on a 16GB card — the level below the 70B-class reasoners most guides assume.

24.4K HF downloads305 likesmistralai/Magistral-Small-2509· stats from 8/14/2026
Consumer GPUMac / Apple Silicon

131K

Max Context

5

Quant Variants

GGUF Q6_K

Best Quality

99.4%

Accuracy Retained

Quantization Variants

Per-quant VRAM, quality loss, and inference speed on RTX 4090

Measured = site benchmarks · Estimated = formula · Community = public reports

FormatLevelBPWVRAMPPL LossSpeedSourceActions
GGUFQ4_K_M4.8514.3 GB2.9%60 tok/sCommunity
CalcHF
GGUFQ5_K_M5.6816.8 GB1.4%54 tok/sEstimated
CalcHF
GGUFQ6_K6.5619.4 GB0.6%47 tok/sEstimated
CalcHF
AWQINT4413.0 GB4.0%76 tok/sEstimated
CalcHF
EXL24.65bpw4.6513.9 GB2.5%86 tok/sEstimated
CalcHF