Qwen3 1.7B Instruct
1.7BAlibaba Qwen3
Tiny Qwen3 with thinking mode. Q4 ~1.4GB — phones, NUC, and always-on local agents.
33K
Max Context
3
Quant Variants
GGUF Q8_0
Best Quality
99.6%
Accuracy Retained
Quantization Variants
Per-quant VRAM, quality loss, and inference speed on RTX 4090
Measured = site benchmarks · Estimated = formula · Community = public reports
Similar models
Compare with Qwen3 8BQwen3 8B Instruct
Alibaba Qwen3
Latest Qwen3 dense 8B with thinking mode. Strong upgrade from Qwen2.5 7B for local deploy.
Qwen3 4B Instruct
Alibaba Qwen3
Smallest Qwen3 dense with thinking mode. Q4 ~3.2GB — ideal for 8GB GPUs and edge devices.
Qwen3 14B Instruct
Alibaba Qwen3
Qwen3 14B — best balance of reasoning and VRAM in the 2026 Qwen lineup.
Qwen3 30B-A3B Instruct
Alibaba Qwen3
Qwen3 MoE with only 3B active params. Q4 ~19GB file; outperforms QwQ-32B on 16GB cards.
How to actually run this
Deployment guides for this model and this class of hardware.