GPU compatibility
Comfortable 14B Q4, some 20B-class models. Examples: RTX 4060 Ti 16GB, RTX 4080, RTX 5070 Ti, RTX 5080.
This preset is useful locally, but model choice matters. Stay close to the green list for the best experience.
14 models fit inside the clean planning capacity.
19 models can run with RAM/offload tradeoffs.
Examples to avoid locally: gpt-oss 120B, gpt-oss-safeguard 120B, Devstral 2 123B.
Small local reasoning, routing, tool decisions, and lightweight coding on 6GB-8GB GPUs
Small multimodal local assistant and low-resource setups
Small local coding assistant and agent tool generation
Fast local chat and simple agent tasks
Local reasoning, agent planning, and tool-use workflows on 16GB+ GPUs
Local safety-policy classification, input-output filtering, moderation review, and trust-and-safety labeling on 16GB+ GPUs
Software engineering agents, repo navigation, patch planning, and local coding workflows
24GB-class multimodal agent, coding assistant, and reasoning workloads