Qwen3 4B Instruct
4BAlibaba Qwen3
Smallest Qwen3 dense with thinking mode. Q4 ~3.2GB — ideal for 8GB GPUs and edge devices.
33K
Max Context
4
Quant Variants
GGUF Q6_K
Best Quality
99.3%
Accuracy Retained
Quantization Variants
Per-quant VRAM, quality loss, and inference speed on RTX 4090
Measured = site benchmarks · Estimated = formula · Community = public reports
Similar models
Compare with Qwen3 8BQwen3 8B Instruct
Alibaba Qwen3
Latest Qwen3 dense 8B with thinking mode. Strong upgrade from Qwen2.5 7B for local deploy.
Qwen3 14B Instruct
Alibaba Qwen3
Qwen3 14B — best balance of reasoning and VRAM in the 2026 Qwen lineup.
Qwen3 30B-A3B Instruct
Alibaba Qwen3
Qwen3 MoE with only 3B active params. Q4 ~19GB file; outperforms QwQ-32B on 16GB cards.
Qwen3 32B Instruct
Alibaba Qwen3
Qwen3 dense 32B — successor to Qwen2.5-32B with stronger reasoning and thinking mode.
How to actually run this
Deployment guides for this model and this class of hardware.