Qwen3-Coder 30B-A3B Instruct
30B-A3BAlibaba Qwen3
Agentic coding MoE with 3.3B active params and 256K native context. Top open coder for 16–24GB cards.
262K
Max Context
3
Quant Variants
GGUF Q5_K_M
Best Quality
99.0%
Accuracy Retained
Quantization Variants
Per-quant VRAM, quality loss, and inference speed on RTX 4090
Measured = site benchmarks · Estimated = formula · Community = public reports
Similar models
Compare with Qwen3 30B-A3BQwen3 30B-A3B Instruct
Alibaba Qwen3
Qwen3 MoE with only 3B active params. Q4 ~19GB file; outperforms QwQ-32B on 16GB cards.
Qwen3 32B Instruct
Alibaba Qwen3
Qwen3 dense 32B — successor to Qwen2.5-32B with stronger reasoning and thinking mode.
Qwen3 8B Instruct
Alibaba Qwen3
Latest Qwen3 dense 8B with thinking mode. Strong upgrade from Qwen2.5 7B for local deploy.
Qwen3 14B Instruct
Alibaba Qwen3
Qwen3 14B — best balance of reasoning and VRAM in the 2026 Qwen lineup.
How to actually run this
Deployment guides for this model and this class of hardware.