2026-09-24 / 9 min read
Which Local LLM Is Most Useful With 36 GB VRAM and 128 GB RAM?
I tested three MoE models from 125B to 321B parameters to find which one gives the best mix of capability and speed on my workstation.
The question
How much model can this workstation use well?
The test machine
- CPU
- Intel Core i5-14500
- Memory
- 128 GiB DDR5-4000 (4 × 32 GiB; rated 4800 MT/s)
- GPUs
- RTX 3090 24 GiB + RTX 3060 12 GiB
- Serving
- llama.cpp CUDA / Podman
- Context
- 131,072 tokens
The models
Three large mixture-of-experts models, four configurations
- Qwen 3.8 Flash Next: 125B total / 6B active parameters, UD-IQ4_XS, 87.25 GiB.
- DeepSeek V4 Flash: 284B total / 13B active parameters, UD-IQ3_XXS, 97.05 GiB.
- GLM 5.3 Flash (GLM IQ1): 321B total / 18B active parameters, UD-IQ1_S, 86.69 GiB.
- GLM 5.3 Flash (GLM Q3): the same model at UD-Q3_K_XL, 137.40 GiB.
Method
Coding, reasoning, and serving speed
The GLM Q3 run a week later used the same harness, sampled questions, 128K context, output limits, and greedy decoding. It added 374 responses and 5.2 hours.
Measured comparison
Qwen led HumanEval; GLM Q3 edged ahead on MMLU-Pro
Exact-match accuracy. Each benchmark places all four deployable model configurations on the same zero-based scale.
HumanEval
164 problemsMMLU-Pro
210 questionsResult
The biggest model was not the strongest model
GLM Q3 reached 89.02% on HumanEval and 84.76% on MMLU-Pro. That is the best MMLU-Pro score in this test, but only two questions ahead of Qwen: 178/210 against 176/210.
The result I did not expect
125B beat 321B at almost the same memory footprint
Parameter counts are difficult to compare across mixture-of-experts architectures. For each token, the model uses only a selection of its experts. Their design, routing, training, and tuning can matter more than the number of active parameters.
There is also a much simpler explanation in the files I tested. Qwen’s UD-IQ4_XS model takes 87.25 GiB. GLM fits into 86.69 GiB only because it uses the much more aggressive UD-IQ1_S quantization. Both deployments demand almost the same amount of memory, but Qwen keeps more information per parameter. I count that difference as part of the result because choosing a quantization that fits is part of choosing a local model.
The paired results make the difference easier to see. Both models solved 140 HumanEval problems. Qwen alone solved 18, GLM alone solved two, and both missed four. GLM also averaged 258 output tokens per answer, compared with 81 for Qwen. Six GLM answers reached the output limit, and every one failed. Concise, well-formed output is part of being useful, even when the underlying problem is excessive generation rather than reasoning.
On this machine, this Qwen configuration is the better system. The test does not tell us whether the unquantized Qwen model is generally smarter than unquantized GLM. The sampled MMLU-Pro scores were also much closer: 176/210 for Qwen and 168/210 for GLM, with overlapping confidence intervals.
Update
The quantization explained part of the gap, not all of it
On MMLU-Pro, GLM Q3 solved 178 questions, ten more than GLM IQ1 and two more than Qwen. Its confidence interval still overlaps Qwen’s.
So the 1-bit quantization cost GLM real quality, but a much larger GLM file still did not catch Qwen at coding. The price is memory: GLM Q3 needs 137 GiB against Qwen’s 87 GiB, which uses almost everything this machine has.
Cost in time
Qwen was also the fastest to evaluate
Sequential evaluation time, excluding model loading. Lower is better.
HumanEval
hoursMMLU-Pro
hoursServing speed
Three of four configurations cleared 10 tok/s
Generation throughput reported by llama.cpp during MMLU-Pro. Higher is better.
MMLU-Pro generation
10 tok/s targetConclusion
Yes, within limits
I would use these models for drafting code, explaining an unfamiliar system, generating tests, and working through a contained problem. HumanEval cannot tell me whether they could maintain a real codebase on their own. It does tell me that I can get useful work from a local model without giving up privacy or a large context window.