Today's Local LLM Pick: qwq:32b on RTX 3090 (2026)
Daily 3090 recommendation for qwq:32b: deliberate performer at 25.3 tok/s, RTX 3090 benchmark data, use-case fit, and local-vs-cloud decision guide.
Fast verdict
qwq:32b runs at 25.3 tok/s on a 24GB RTX 3090 — in the deliberate range. This model prioritizes quality or parameter count over raw speed. Test it on offline or background tasks first, and consider a smaller quantization if interactive response time matters.
qwq:32b approaches the 24GB boundary at higher quantizations. Consider Q4 or Q5 if you need context headroom on the RTX 3090. It ranks #13 of 18 in throughput among currently measured models on this RTX 3090. The next faster model is qwen2.5-coder:32b (31.2 tok/s, 24% faster). The next slower model is glm-4.7-flash:bf16 (11.2 tok/s, 125% slower).
The daily goal is simple: help a 3090 owner decide what to download tonight, what to skip, and when a cloud fallback is the better use of time.
Today’s pick
- Model:
qwq:32b - Category: reasoning
- Size tier: large
- Performance tier: deliberate
- RTX 3090 speed: 25.3 tok/s
- Latency: 4187 ms
- Test time: 2026-07-08T15:58:01Z
- Baseline command:
ollama run qwq:32b
Who should try it
- Users working on math, logic, planning, or multi-step reasoning tasks where
qwq:32b’s chain-of-thought adds accuracy. - Researchers and power users who want a local alternative to cloud reasoning APIs like o1 or Claude.
- Anyone curious whether local reasoning models have caught up to cloud counterparts on a 24GB RTX 3090.
Who should skip it
- Teams that need fast, single-turn responses for real-time applications; reasoning models trade speed for depth.
- Users running simple classification or extraction tasks that don’t benefit from extended reasoning chains.
- Anyone deploying to production without first validating output quality on representative data.
Watch points
- Over-thinking risk: on simple prompts the model may produce unnecessary chain-of-thought, increasing latency.
- Temperature tuning: lower temperatures (0–0.3) improve factual accuracy; higher values may hallucinate reasoning steps.
- Batch efficiency: for throughput-critical tasks, group prompts and process offline rather than requesting real-time responses.
Verified benchmark anchors
gpt-oss:20b: 156.1 tok/s | latency 1524 ms | test 2026-04-29T05:39:58Zqwen3-coder:30b: 149.8 tok/s | latency 897 ms | test 2026-07-08T15:58:01Zqwen3:8b: 112.7 tok/s | latency 1536 ms | test 2026-07-08T15:58:01Zqwen2.5:14b: 84.0 tok/s | latency 946 ms | test 2026-04-29T05:39:58Zministral-3:14b: 80.7 tok/s | latency 2003 ms | test 2026-07-08T15:58:01Z
RTX 3090 decision guide
- Offline first: prioritize qwq:32b for scheduled batch inference, research, or validation workflows.
- Context is the bottleneck: reduce context to the minimum viable length for your task.
- Quantize before you buy hardware: Q4 or Q5 may make this viable on 24GB where Q8 is not.
- Cloud for interactive: if real-time response is required, treat qwq:32b as a cloud-fallback candidate.
Comparisons to validate
qwq:32bvs the next-fastest and next-slowest model in the benchmark feed.qwq:32bvsgpt-oss:20b— same size tier, 25 vs 156 tok/s.qwq:32blocal power cost vs A100 rental for the same workload.
Next actions
- Estimate VRAM fit: /en/tools/vram-calculator/
- Model page: /en/models/qwq-32b-q4/
- Benchmark changelog: /en/benchmarks/changelog/
- Local hardware path: /en/affiliate/hardware-upgrade/
- Cloud fallback: /go/runpod and /go/vast
Affiliate Disclosure: This post may include affiliate links. LocalVRAM may earn a commission at no extra cost.
Related model pages
Model profiles referenced by this article, with VRAM fit and measured or estimated status.
- QwQ 32B Q4 qwq:32b · Q4 · 18-28GB VRAM · estimated
- Ministral 3 14B Q4 ministral-3:14b · Q4 · 12-14GB VRAM · measured
- Glm 4.7 Flash 7B Q4 glm-4.7-flash:bf16 · Q4 · 6-16GB VRAM · measured
- GPT-OSS 20B Q4 gpt-oss:20b · Q4 · 16-20GB VRAM · estimated
- Qwen3 8B Q4 qwen3:8b · Q4 · 8-10GB VRAM · estimated