Local LLM model fit
Qwen3.8-Flash-Next 125B-A6B is a 125B Qwen3.8 model. This page estimates Q4 VRAM fit, Ollama command, context planning, and fallback choices for common local AI GPUs.
Experimental local Qwen4-architecture testing, multimodal coding agents, and long-context work on very large-memory systems. Weakness: The official Q4_K_M Ollama file alone is 120GB; the 144GB runtime figure is a planning estimate, not a measured benchmark.
| Hardware | Examples | Clean capacity | Q4 need | Status | Calculator |
|---|---|---|---|---|---|
| 6 GB VRAM entry GPU | GTX 1660, RTX 2060 6GB | 4.5 GB usable VRAM | 144 GB | Too large | Open calculator |
| 8 GB VRAM mainstream GPU | RTX 3060 Ti, RTX 4060, RTX 3070 | 6.5 GB usable VRAM | 144 GB | Too large | Open calculator |
| 10 GB VRAM older high-end GPU | RTX 3080 10GB | 8.5 GB usable VRAM | 144 GB | Too large | Open calculator |
| 12 GB VRAM local agent GPU | RTX 3060 12GB, RTX 4070, RTX 5070 | 10.5 GB usable VRAM | 144 GB | Too large | Open calculator |
| 16 GB VRAM creator GPU | RTX 4060 Ti 16GB, RTX 4080, RTX 5070 Ti, RTX 5080 | 14.5 GB usable VRAM | 144 GB | Too large | Open calculator |
| 24 GB VRAM homelab workstation | RTX 3090, RTX 4090 | 22.5 GB usable VRAM | 144 GB | Too large | Open calculator |
| 32 GB VRAM Blackwell workstation | RTX 5090 | 30.5 GB usable VRAM | 144 GB | Too large | Open calculator |
| 48 GB VRAM workstation | RTX A6000, L40S 48GB | 46.5 GB usable VRAM | 144 GB | Too large | Open calculator |
| Apple Silicon 32 GB unified memory | M2 Max 32GB, M3 Max 36GB | 26 GB unified | 144 GB | Too large | Open calculator |
| Apple Silicon 256 GB unified memory | Mac Studio M3 Ultra 256GB, Mac Studio M4 Ultra 256GB | 250 GB unified | 144 GB | Runs locally | Open calculator |
| Quantization | Estimated memory | Use case |
|---|---|---|
| Q4 / 4-bit | 144 GB | Default local inference balance |
This is a practical planning estimate, not a benchmark. Real memory use changes with backend, context length, KV cache, quantization file, drivers, and offloading settings.
123B · High-end local software engineering agents, large-repository navigation, and tool-heavy coding workflows on server-class memory
120B · Large local reasoning servers, heavy agent orchestration, and high-end homelab inference
117B · High-end local safety reasoning, nuanced policy decisions, offline labeling, and moderation pipelines on H100-class or 80GB+ systems
80B · High-end local coding agents, repository-scale code edits, and tool-calling development workflows