Back to Quant Hub

Phi-4 Mini Instruct

3.8B

Microsoft Phi

Latest Phi mini with improved math and code. Strong 4B-class performer.

Consumer GPUMac / Apple SiliconCPU / VPS

131K

Max Context

3

Quant Variants

GGUF Q8_0

Best Quality

99.8%

Accuracy Retained

Quantization Variants

Per-quant VRAM, quality loss, and inference speed on RTX 4090

FormatLevelBPWVRAMPPL LossSpeedActions
GGUFQ4_K_M4.852.8 GB3.5%305 tok/s
GGUFQ8_08.54.4 GB0.2%262 tok/s
AWQINT442.5 GB4.8%395 tok/s