Qwen3-VL 30B-A3B Instruct
30B-A3BAlibaba Qwen3-VL
Multimodal MoE with only ~3B active parameters, so it stays responsive on Apple unified memory and survives CPU offload far better than a dense 30B. Q4 ~19GB — a 24GB card holds it outright.
41K
Max Context
4
Quant Variants
GGUF Q8_0
Best Quality
99.7%
Accuracy Retained
Quantization Variants
Per-quant VRAM, quality loss, and inference speed on RTX 4090
Measured = site benchmarks · Estimated = formula · Community = public reports
Similar models
Compare with Qwen3-VL 8BQwen3-VL 8B Instruct
Alibaba Qwen3-VL
Current-generation vision-language model that still fits a single 8–12GB card at Q4 (~5.9GB). The realistic multimodal option for people without a 24GB GPU — note the vision encoder adds VRAM the KV-cache math below does not model.
Gemma 3 27B IT
Google Gemma 3
Gemma 3 large instruct with long context and multimodal support. Q4 ~16GB — dual-GPU or 24GB card with short ctx.
Magistral Small 1.2 24B
Mistral Magistral
Mistral's reasoning model on a Mistral Small 3.2 base, with reasoning wrapped in [THINK] tags. Q4 ~14GB puts explicit chain-of-thought on a 16GB card — the level below the 70B-class reasoners most guides assume.
Seed-OSS 36B Instruct
ByteDance Seed
Dense 36B with a native 512K context. Q4 weights are ~22GB, which lands on a 24GB card or a 2×16GB split — but the context is the real cost: 512K of KV cache is roughly 128GB on its own, so budget context first and weights second.
How to actually run this
Deployment guides for this model and this class of hardware.