Embedding Models

14 profiles in this group. Use this hub page to compare practical VRAM floor, expected throughput, and best local-vs-cloud path.

RAM planning is included because Ollama may spill layers to CPU when VRAM is tight, especially with long context.

Model Data VRAM min VRAM optimal System RAM RTX 3090 fit Best local GPU Cloud fallback Detail
Qwen3 Embedding 8B FP16 Estimated 18GB 30GB 64GB+ Borderline RTX 6000 Ada 48GB A100 80GB Open
Qwen3 Embedding 8B Q4 Estimated 6GB 16GB 32GB+ Comfortable RTX 3090 24GB A6000 48GB Open
Qwen3 Embedding 8B Q5 Estimated 8GB 18GB 32GB+ Comfortable RTX 3090 24GB A6000 48GB Open
Qwen3 Embedding 8B Q8 Estimated 12GB 22GB 32GB+ Comfortable RTX 3090 24GB A6000 48GB Open
BGE-M3 567M FP16 Estimated 4GB 12GB 16GB+ Comfortable RTX 3090 24GB A6000 48GB Open
MXBAI Embed Large 335M FP16 Estimated 2GB 10GB 16GB+ Comfortable RTX 3090 24GB A6000 48GB Open
Snowflake Arctic Embed 335M FP16 Estimated 2GB 10GB 16GB+ Comfortable RTX 3090 24GB A6000 48GB Open
Nomic Embed Text 137M FP16 Estimated 2GB 10GB 16GB+ Comfortable RTX 3090 24GB A6000 48GB Open
Snowflake Arctic Embed 137M FP16 Estimated 2GB 10GB 16GB+ Comfortable RTX 3090 24GB A6000 48GB Open
Snowflake Arctic Embed 110M FP16 Estimated 2GB 10GB 16GB+ Comfortable RTX 3090 24GB A6000 48GB Open
All-MiniLM 33M FP16 Estimated 2GB 10GB 16GB+ Comfortable RTX 3090 24GB A6000 48GB Open
Snowflake Arctic Embed 33M FP16 Estimated 2GB 10GB 16GB+ Comfortable RTX 3090 24GB A6000 48GB Open
All-MiniLM 22M FP16 Estimated 2GB 10GB 16GB+ Comfortable RTX 3090 24GB A6000 48GB Open
Snowflake Arctic Embed 22M FP16 Estimated 2GB 10GB 16GB+ Comfortable RTX 3090 24GB A6000 48GB Open
Back to all groups Use VRAM calculator Run large models on RunPod Try Vast.ai fallback

We may earn a commission if you click links on this page.