GPU compatibility
Heavy local models and homelab inference. Examples: RTX 3090, RTX 4090.
This preset has enough clean Q4 headroom for most curated local models in this dataset.
23 models fit inside the clean planning capacity.
10 models can run with RAM/offload tradeoffs.
Examples to avoid locally: gpt-oss 120B, gpt-oss-safeguard 120B, Devstral 2 123B.
Small local reasoning, routing, tool decisions, and lightweight coding on 6GB-8GB GPUs
Small multimodal local assistant and low-resource setups
Small local coding assistant and agent tool generation
Fast local chat and simple agent tasks
High-quality multimodal reasoning, coding assistants, and local-first agent workflows
Higher-quality local multimodal reasoning, screenshot analysis, document/image extraction, and GUI-agent planning
Strong local coding and architecture work on 24GB GPUs
Heavy local reasoning on 24GB GPUs