GPU compatibility
Local routing, agents, and model testing. Examples: RTX 3060 12GB, RTX 4070, RTX 5070.
This preset is useful locally, but model choice matters. Stay close to the green list for the best experience.
9 models fit inside the clean planning capacity.
21 models can run with RAM/offload tradeoffs.
Examples to avoid locally: Llama 3.1 70B Instruct, Qwen3-Next 80B-A3B Instruct, Qwen3-Coder-Next.
Small local reasoning, routing, tool decisions, and lightweight coding on 6GB-8GB GPUs
Small multimodal local assistant and low-resource setups
Small local coding assistant and agent tool generation
Fast local chat and simple agent tasks
Balanced multimodal local chat on 12GB+ GPUs
Efficient multimodal local assistant and edge-style agent workflows
Local coding, scripts, repo assistance, technical agents
Local reasoning and debugging on 12GB/16GB GPUs