Gemma 3 27B IT
27BGoogle Gemma 3
Gemma 3 large instruct with long context and multimodal support. Q4 ~16GB — dual-GPU or 24GB card with short ctx.
131K
Max Context
3
Quant Variants
GGUF Q4_K_M
Best Quality
97.2%
Accuracy Retained
Quantization Variants
Per-quant VRAM, quality loss, and inference speed on RTX 4090
Measured = site benchmarks · Estimated = formula · Community = public reports
Similar models
Compare with Gemma 3Gemma 3 4B IT
Google Gemma 3
Google Gemma 3 multimodal 4B. 128K context; strong vision + text on 8GB cards.
Gemma 3 12B IT
Google Gemma 3
Mid-size Gemma 3 with vision. Fits 16GB at Q4; excellent multilingual chat.
Qwen3-VL 30B-A3B Instruct
Alibaba Qwen3-VL
Multimodal MoE with only ~3B active parameters, so it stays responsive on Apple unified memory and survives CPU offload far better than a dense 30B. Q4 ~19GB — a 24GB card holds it outright.
Magistral Small 1.2 24B
Mistral Magistral
Mistral's reasoning model on a Mistral Small 3.2 base, with reasoning wrapped in [THINK] tags. Q4 ~14GB puts explicit chain-of-thought on a 16GB card — the level below the 70B-class reasoners most guides assume.
How to actually run this
Deployment guides for this model and this class of hardware.