Gemma 3 4B IT
4BGoogle Gemma 3
Google Gemma 3 multimodal 4B. 128K context; strong vision + text on 8GB cards.
131K
Max Context
3
Quant Variants
GGUF Q8_0
Best Quality
99.8%
Accuracy Retained
Quantization Variants
Per-quant VRAM, quality loss, and inference speed on RTX 4090
Measured = site benchmarks · Estimated = formula · Community = public reports
Similar models
Compare with Gemma 3Gemma 3 12B IT
Google Gemma 3
Mid-size Gemma 3 with vision. Fits 16GB at Q4; excellent multilingual chat.
Gemma 3 27B IT
Google Gemma 3
Gemma 3 large instruct with long context and multimodal support. Q4 ~16GB — dual-GPU or 24GB card with short ctx.
Llama 3.1 8B Instruct
Meta Llama 3.1
Meta's flagship 8B model with 128K context. Best-in-class for local deployment.
Qwen2.5 7B Instruct
Alibaba Qwen2.5
Alibaba's highly optimized 7B. Punches well above its weight, especially in coding.
How to actually run this
Deployment guides for this model and this class of hardware.