GPU compatibility
Unified memory; not directly comparable to discrete VRAM. Examples: M2 Max 32GB, M3 Max 36GB.
This preset has enough clean Q4 headroom for most curated local models in this dataset.
27 models fit inside the clean planning capacity.
2 models can run with RAM/offload tradeoffs.
Examples to avoid locally: Mixtral 8x7B, Llama 3.1 70B Instruct, Qwen3-Next 80B-A3B Instruct.
Small local reasoning, routing, tool decisions, and lightweight coding on 6GB-8GB GPUs
Small multimodal local assistant and low-resource setups
Small local coding assistant and agent tool generation
Fast local chat and simple agent tasks
Current 27B-class local coding, multimodal analysis, and long-running tool agents on 24GB-class GPUs
Agentic coding and multimodal reasoning when 27B is not enough and 32GB-class headroom is available