Best Local RAG Models for Ollama in 2026
Strong local RAG results come from pairing a fast embedding model with a generation model that fits your VRAM. Use this page to choose a 16GB/24GB-friendly stack before moving to cloud GPUs.
Embedding models for local retrieval quality
| Model | VRAM min/optimal | Category | Detail |
|---|---|---|---|
| MXBAI Embed Large 335M FP16 | 2GB / 10GB | embedding | Open |
| Snowflake Arctic Embed 335M FP16 | 2GB / 10GB | embedding | Open |
| Nomic Embed Text 137M FP16 | 2GB / 10GB | embedding | Open |
| Snowflake Arctic Embed 137M FP16 | 2GB / 10GB | embedding | Open |
| Snowflake Arctic Embed 110M FP16 | 2GB / 10GB | embedding | Open |
| All-MiniLM 33M FP16 | 2GB / 10GB | embedding | Open |
| Snowflake Arctic Embed 33M FP16 | 2GB / 10GB | embedding | Open |
| All-MiniLM 22M FP16 | 2GB / 10GB | embedding | Open |
| Snowflake Arctic Embed 22M FP16 | 2GB / 10GB | embedding | Open |
| BGE-M3 567M FP16 | 4GB / 12GB | embedding | Open |
Generation models that fit local VRAM
| Model | VRAM min/optimal | 3090 tok/s | Data | Detail |
|---|---|---|---|---|
| Qwen 0.5B Q4 | 2GB / 10GB | 48 | Estimated | Open |
| Qwen2 0.5B Q4 | 2GB / 10GB | 48 | Estimated | Open |
| SmolLM2 360M Q4 | 2GB / 10GB | 48 | Estimated | Open |
| SmolLM2 135M Q4 | 2GB / 10GB | 48 | Estimated | Open |
| Gemma 2 2B Q4 | 2GB / 12GB | 42 | Estimated | Open |
| Gemma 2B Q4 | 2GB / 12GB | 42 | Estimated | Open |
| Gemma4 E2B Q4 | 2GB / 12GB | 42 | Estimated | Open |
| Qwen 1.8B Q4 | 2GB / 12GB | 42 | Estimated | Open |
| SmolLM2 1.7B Q4 | 2GB / 12GB | 42 | Estimated | Open |
| DeepSeek-R1 1.5B Q4 | 2GB / 12GB | 42 | Estimated | Open |
| Qwen2 1.5B Q4 | 2GB / 12GB | 42 | Estimated | Open |
| TinyLlama 1.1B Q4 | 2GB / 12GB | 42 | Estimated | Open |
Recommended local RAG flow
- Pick an embedding model that fits your latency and memory budget.
- Use a 7B to 32B generator model that can sustain your expected context length.
- Tune retrieval quality first, then upgrade generation model size if needed.