GPU compatibility
Ultra high-end consumer workstation for 30B+ models with extra headroom. Examples: RTX 5090.
This preset has enough clean Q4 headroom for most curated local models in this dataset.
29 models fit inside the clean planning capacity.
6 models can run with RAM/offload tradeoffs.
Examples to avoid locally: GLM-5.2.
Small local reasoning, routing, tool decisions, and lightweight coding on 6GB-8GB GPUs
Small multimodal local assistant and low-resource setups
Small local coding assistant and agent tool generation
Fast local chat and simple agent tasks
High-quality local chat and reasoning on workstation-class hardware
High-end local reasoning, long-context planning, and tool-agent workloads when 48GB+ memory is available
High-end local coding agents, repository-scale code edits, and tool-calling development workflows
Large local reasoning servers, heavy agent orchestration, and high-end homelab inference