Tomas Kral Connect

2026-09-24 / 9 min read

Which Local LLM Is Most Useful With 36 GB VRAM and 128 GB RAM?

I tested three MoE models from 125B to 321B parameters to find which one gives the best mix of capability and speed on my workstation.

The question

How much model can this workstation use well?

I wanted a local model with a 128K context window that could generate at around 10 tokens per second and still handle real coding and reasoning work. All three models could run on the machine. This test was about whether they remained useful once loaded.

The test machine

CPU
Intel Core i5-14500
Memory
128 GiB DDR5-4000 (4 × 32 GiB; rated 4800 MT/s)
GPUs
RTX 3090 24 GiB + RTX 3060 12 GiB
Serving
llama.cpp CUDA / Podman
Context
131,072 tokens

The models

Three large mixture-of-experts models, four configurations

  • Qwen 3.8 Flash Next: 125B total / 6B active parameters, UD-IQ4_XS, 87.25 GiB.
  • DeepSeek V4 Flash: 284B total / 13B active parameters, UD-IQ3_XXS, 97.05 GiB.
  • GLM 5.3 Flash (GLM IQ1): 321B total / 18B active parameters, UD-IQ1_S, 86.69 GiB.
  • GLM 5.3 Flash (GLM Q3): the same model at UD-Q3_K_XL, 137.40 GiB.
Each score belongs to that exact model and quantization on my machine. I did not try to calculate how much quality each quantization removed. I care about the version I can actually run.

Method

Coding, reasoning, and serving speed

I ran the complete 164-problem HumanEval set for code generation and a 210-question sample of MMLU-Pro for general reasoning, with 15 questions from each of 14 subjects. Models were served through llama.cpp and evaluated sequentially with EleutherAI's lm-evaluation-harness. The run produced 1,122 evaluated responses and took 14.7 hours wall time.

The GLM Q3 run a week later used the same harness, sampled questions, 128K context, output limits, and greedy decoding. It added 374 responses and 5.2 hours.

Measured comparison

Qwen led HumanEval; GLM Q3 edged ahead on MMLU-Pro

Exact-match accuracy. Each benchmark places all four deployable model configurations on the same zero-based scale.

HumanEval

164 problems
Qwen 96.34%
DeepSeek 91.46%
GLM IQ1 86.59%
GLM Q3 89.02%

MMLU-Pro

210 questions
Qwen 83.81%
DeepSeek 80.00%
GLM IQ1 80.00%
GLM Q3 84.76%

Result

The biggest model was not the strongest model

Among the original three, Qwen finished first on HumanEval at 96.34% and on MMLU-Pro at 83.81%. DeepSeek reached 91.46% and 80.00%; GLM IQ1 reached 86.59% and 80.00%. Qwen's 3.81-point MMLU-Pro lead looks promising, although 210 questions per model are too few to call it a clear win.

GLM Q3 reached 89.02% on HumanEval and 84.76% on MMLU-Pro. That is the best MMLU-Pro score in this test, but only two questions ahead of Qwen: 178/210 against 176/210.

The result I did not expect

125B beat 321B at almost the same memory footprint

Qwen has less than half of GLM's total parameters and only a third as many active parameters. I expected GLM's extra capacity to help. Instead, Qwen solved 158 HumanEval problems while GLM solved 142.

Parameter counts are difficult to compare across mixture-of-experts architectures. For each token, the model uses only a selection of its experts. Their design, routing, training, and tuning can matter more than the number of active parameters.

There is also a much simpler explanation in the files I tested. Qwen’s UD-IQ4_XS model takes 87.25 GiB. GLM fits into 86.69 GiB only because it uses the much more aggressive UD-IQ1_S quantization. Both deployments demand almost the same amount of memory, but Qwen keeps more information per parameter. I count that difference as part of the result because choosing a quantization that fits is part of choosing a local model.

The paired results make the difference easier to see. Both models solved 140 HumanEval problems. Qwen alone solved 18, GLM alone solved two, and both missed four. GLM also averaged 258 output tokens per answer, compared with 81 for Qwen. Six GLM answers reached the output limit, and every one failed. Concise, well-formed output is part of being useful, even when the underlying problem is excessive generation rather than reasoning.

On this machine, this Qwen configuration is the better system. The test does not tell us whether the unquantized Qwen model is generally smarter than unquantized GLM. The sampled MMLU-Pro scores were also much closer: 176/210 for Qwen and 168/210 for GLM, with overlapping confidence intervals.

Update

The quantization explained part of the gap, not all of it

Running GLM at UD-Q3_K_XL tested the simpler explanation directly. The coding gap to Qwen narrowed from 16 problems to 12: GLM Q3 solved 146 HumanEval problems. Both models solved 145, Qwen alone solved 13, GLM Q3 alone solved one, and both missed five. GLM Q3's answers were also about a third shorter than GLM IQ1's, and only three reached the output limit instead of six.

On MMLU-Pro, GLM Q3 solved 178 questions, ten more than GLM IQ1 and two more than Qwen. Its confidence interval still overlaps Qwen’s.

So the 1-bit quantization cost GLM real quality, but a much larger GLM file still did not catch Qwen at coding. The price is memory: GLM Q3 needs 137 GiB against Qwen’s 87 GiB, which uses almost everything this machine has.

Cost in time

Qwen was also the fastest to evaluate

Sequential evaluation time, excluding model loading. Lower is better.

HumanEval

hours
Qwen 0.28 h
DeepSeek 0.52 h
GLM IQ1 1.11 h
GLM Q3 1.17 h

MMLU-Pro

hours
Qwen 2.79 h
DeepSeek 5.90 h
GLM IQ1 3.84 h
GLM Q3 3.99 h

Serving speed

Three of four configurations cleared 10 tok/s

Generation throughput reported by llama.cpp during MMLU-Pro. Higher is better.

MMLU-Pro generation

10 tok/s target
Qwen 18.8
DeepSeek 9.53
GLM IQ1 12.08
GLM Q3 10.18

Conclusion

Yes, within limits

Qwen was the practical winner: the highest HumanEval score, an MMLU-Pro score within two questions of the best, 18.8 tok/s during MMLU-Pro, and the shortest total evaluation time. GLM proved that a 321B-parameter MoE model can still clear 10 tok/s on this split CPU/GPU setup, although its coding score and runtime were weaker. At UD-Q3_K_XL it became the strongest reasoner in the test, but it needs 50 GiB more memory than Qwen and runs at 10.18 tok/s, right at my target. DeepSeek remained capable, but its measured 9.53 tok/s narrowly missed my speed target.

I would use these models for drafting code, explaining an unfamiliar system, generating tests, and working through a contained problem. HumanEval cannot tell me whether they could maintain a real codebase on their own. It does tell me that I can get useful work from a local model without giving up privacy or a large context window.