Qwen3 8B Instruct
8BAlibaba Qwen3
Latest Qwen3 dense 8B with thinking mode. Strong upgrade from Qwen2.5 7B for local deploy.
41K
Max Context
4
Quant Variants
GGUF Q6_K
Best Quality
99.4%
Accuracy Retained
Quantization Variants
Per-quant VRAM, quality loss, and inference speed on RTX 4090
Measured = site benchmarks · Estimated = formula · Community = public reports
Similar models
Compare with Qwen3 4BQwen3 4B Instruct
Alibaba Qwen3
Smallest Qwen3 dense with thinking mode. Q4 ~3.2GB — ideal for 8GB GPUs and edge devices.
Qwen3 14B Instruct
Alibaba Qwen3
Qwen3 14B — best balance of reasoning and VRAM in the 2026 Qwen lineup.
Qwen3 30B-A3B Instruct
Alibaba Qwen3
Qwen3 MoE with only 3B active params. Q4 ~19GB file; outperforms QwQ-32B on 16GB cards.
Qwen3 32B Instruct
Alibaba Qwen3
Qwen3 dense 32B — successor to Qwen2.5-32B with stronger reasoning and thinking mode.
How to actually run this
Deployment guides for this model and this class of hardware.