8 local AI models that run great on 8GB of VRAM or less
www.makeuseof.com/8-local-ai-models-that-run-great-on-8gb-of-vram-or-less/
Test bench and performance expectations. All-AMD system with configurable VRAM limits
- Phi-3.5 Mini (3.8 B). Fast and efficient
- Llama 3.1 (8B). Good general-purpose model
- Mistral 7B (v0.3). Speedy but consistent
- Qwen 2.5 (7B). Great but a bit too formal
- Gemma 2 (9B). The most powerful one
- DeepSeek-R1-Distill-Qwen (7B). DeepSeek for 8GB VRAM
- DeepSeek-Coder (6.7B). Coding specialized DeepSeek
- Gemma 2 (2B). Deep world knowledge will be an issue
- Benchmarking the models. Drawing some conclusions