Gemma 3 12B IT
12BGoogle Gemma 3
Mid-size Gemma 3 with vision. Fits 16GB at Q4; excellent multilingual chat.
131K
Max Context
3
Quant Variants
GGUF Q5_K_M
Best Quality
98.7%
Accuracy Retained
Quantization Variants
Per-quant VRAM, quality loss, and inference speed on RTX 4090
Measured = site benchmarks · Estimated = formula · Community = public reports
Similar models
Compare with Gemma 3Gemma 3 4B IT
Google Gemma 3
Google Gemma 3 multimodal 4B. 128K context; strong vision + text on 8GB cards.
Gemma 3 27B IT
Google Gemma 3
Gemma 3 large instruct with long context and multimodal support. Q4 ~16GB — dual-GPU or 24GB card with short ctx.
Qwen3-VL 8B Instruct
Alibaba Qwen3-VL
Current-generation vision-language model that still fits a single 8–12GB card at Q4 (~5.9GB). The realistic multimodal option for people without a 24GB GPU — note the vision encoder adds VRAM the KV-cache math below does not model.
Qwen2.5 14B Instruct
Alibaba Qwen2.5
The sweet spot between performance and resource usage. 16GB VRAM with Q4.
How to actually run this
Deployment guides for this model and this class of hardware.